Pull down to go back
Transformer Model Optimization: Beyond FP16 + ONNX—Why Pruning & Graph Optimization Aren't Delivering

Transformer Model Optimization: Beyond FP16 + ONNX—Why Pruning & Graph Optimization Aren't Delivering

Transformer 模型優化卡關:FP16 + ONNX 之後還能怎麼做?(剪枝和圖優化效果不大)

A developer shares their struggle optimizing a transformer model for speed and size. They've already halved the model with FP16 conversion, tried ONNX Runtime optimization, and experimented with both unstructured and structured pruning plus graph optimizations—but hit a wall. The model still sits at ~162 MB with no major improvements from recent attempts. They're looking for what actually works when the standard tricks plateau.