Pull down to go back
I benchmarked 21 local LLMs on a MacBook Air M5 for code quality AND speed

I benchmarked 21 local LLMs on a MacBook Air M5 for code quality AND speed

我在 MacBook Air M5 上實測了 21 個本地大型語言模型的程式碼品質和速度

There are plenty of "bro trust me, this model is better for coding" discussions out there. I wanted to replace the vibes with actual data: which model writes correct code and how fast does it run on real hardware, tested under identical conditions so the results are directly comparable. No cherry-picked prompts, no subjective impressions, just pass@1 on 164 coding problems with an expanded test suite. Hardware: MacBook Air M5, 32 GB unified memory Quantization: Q4_K_M for all models via llama.cpp