Pull down to go back
Qwen3.6 Crashes When Using Turboquant—Here's What Happened

Qwen3.6 Crashes When Using Turboquant—Here's What Happened

Qwen3.6 不吃 Turboquant——量化方法大亂鬥

A user tested Qwen3.6 across three different quantization methods (llama.cpp, ik_llama.cpp, and Turboquant) on dual GPUs (3080 + 3060 12GB each) with identical settings. The result? Turboquant caused the model to fail while the other methods ran fine. They're using the TheTom/llama-cpp-turboquant implementation and kept everything constant except the quantization approach (turbo3 vs q8_0). If you're trying to run Qwen3.6 locally on limited VRAM, this is a heads-up that Turboquant might not be the move—at least not yet.