Pull down to go back
Final Monster: 32x AMD MI50 32GB Hits 9.7 t/s Output & 264 t/s Input with Kimi K2.6

Final Monster: 32x AMD MI50 32GB Hits 9.7 t/s Output & 264 t/s Input with Kimi K2.6

終極怪獸:32 顆 AMD MI50 32GB 搭配 Kimi K2.6 達成 9.7 t/s(輸出)& 264 t/s(輸入)

Someone just built an absolute beast: 32 AMD MI50 GPUs running Kimi K2.6 (an open-source LLM) and achieved 9.7 tokens/second for output and 264 tokens/second for input processing. The setup uses a custom vLLM fork optimized for AMD GPUs, pulling about 640W idle and 4.8kW at peak—basically a small power plant. The honest take? It's cool as hell but only makes financial sense if you have free energy (solar, data center waste heat, etc.). It's two 16-GPU nodes connected via 10G ethernet, proving you can DIY a massive inference cluster if you're willing to deal with the power bill.