Pull down to go back
Gaming Mini PC + eGPU Finally Hits Usable Speed: 25.9 tokens/sec with Radeon 9060 XT

Gaming Mini PC + eGPU Finally Hits Usable Speed: 25.9 tokens/sec with Radeon 9060 XT

遊戲迷你電腦配外接顯卡終於跑出實用速度:Radeon 9060 XT 達到 25.9 t/s

So you want to run powerful AI models locally without paying OpenAI every month? I've been experimenting with exactly that—running large language models (LLMs) on a gaming mini PC (AMD 7840HS, 32GB RAM) hooked up to an external GPU (Radeon 9060XT, 16GB VRAM). Honestly, llama.cpp was confusing at first and I kept getting sluggish results, but then I tried the new Gemma4 24B A4B IQ4 NL model and boom—25.9 tokens per second. That's actually fast enough to use. I even plugged it into OpenCode to ask questions about my own codebase, and it actually works. If you've been curious about local AI but thought you needed a $3000 setup, this might change your mind.