Pull down to go back
NVIDIA's Nemotron-3-Super Pruned to Half Size, Fine-Tuned for Math—Now Solves 90%+ of AIME 2026 Problems on a Single GPU

NVIDIA's Nemotron-3-Super Pruned to Half Size, Fine-Tuned for Math—Now Solves 90%+ of AIME 2026 Problems on a Single GPU

NVIDIA Nemotron-3-Super 被砍掉一半大小,微調後數學能力反而更強——單張 GPU 就能解決 90% 以上的 AIME 2026 題目

A developer just released a heavily optimized version of NVIDIA's Nemotron-3-Super-120B that cuts the model size in half (from 120B to 64B parameters) while actually getting *better* at math. Here's what they did: took the original model, pruned its expert system from 512 down to 256 (basically removing redundant parts), fine-tuned it on ~270 competition math problems using a technique called GRPO, then compressed it with AWQ and FP8 quantization. The result? It now scores 90%+ on AIME 2026 (a notoriously hard math competition benchmark) and fits on a single H100 GPU. This is the kind of thing that matters if you're running local AI—smaller model, same or better performance, way cheaper to run. The benchmarks are solid, and they've released multiple versions (BF16 full weights, quantized versions) so you can pick what fits your hardware.