Pull down to go back
I tested Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.5-27B and Gemma 4 on the same real architecture-writing task on an RTX 5090

I tested Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.5-27B and Gemma 4 on the same real architecture-writing task on an RTX 5090

我用 RTX 5090 在同一個真實架構設計任務上測試了 Qwen3.6-27B、Qwen3.6-35B-A3B、Qwen3.5-27B 和 Gemma 4

I ran a straightforward but revealing local-LLM test comparing four models on a real-world architecture-writing task. The test included Gemma4, Qwen3.6-35B, Qwen3.5-27B, and the newly-released Qwen3.6-27B—all quantized versions optimized to run on consumer hardware. I was planning to test just the Qwen models and Gemma4, but Qwen3.6-27B dropped right as I was about to wrap up, so I added it to the comparison. The task involved processing noisy evidence and generating structured architectural decisions, which is a good stress-test for model reasoning and instruction-following. Results show interesting trade-offs between model size, quantization method, and actual performance on this specific task.