Pull down to go back
Running llama.cpp on Snapdragon Hexagon NPU seems promising

Running llama.cpp on Snapdragon Hexagon NPU seems promising

在高通 Snapdragon Hexagon NPU 上執行 llama.cpp 看起來很有潛力

A developer successfully ran llama.cpp on a OnePlus 12 with Snapdragon 8 Gen 3 using the Hexagon NPU backend. Following Qualcomm's official documentation, they cross-compiled the software on Ubuntu and deployed it via Termux. The results are impressive: Gemma-3-12B achieved 8 tokens/sec (prompt processing) and 4.5 tokens/sec (text generation), while the smaller Gemma-3-4B model hit 20 tokens/sec and 12.5 tokens/sec respectively. Qualcomm appears heavily invested in this backend with multiple pull requests from their engineers, suggesting strong official support for running large language models directly on mobile hardware.