Pull down to go back
From 1 token/sec to 100 tokens/sec: Your old hardware just became a supercomputer

From 1 token/sec to 100 tokens/sec: Your old hardware just became a supercomputer

從每秒1個token到100個token:你的舊硬體突然變成超級電腦了

Remember when running Llama 405B at 1.2 tokens/sec locally felt like a huge win? Yeah, that was just two years ago. Now that same hardware can run cutting-edge models like Qwen 3.5-397B, DeepSeek V4 Flash, and Minimax 2.7 at 30-100 tokens/sec—and they absolutely demolish Llama 405B's performance. The speed gains are insane. What used to require a fortune in GPUs now runs on equipment people already own. This is the kind of progress that makes you wonder what's even possible next year.