Pull down to go back
A Simple Training Tweak Makes AI Models 63% Better at Human Judgment—Without Changing Performance Metrics

A Simple Training Tweak Makes AI Models 63% Better at Human Judgment—Without Changing Performance Metrics

訓練方法小改變,AI 模型在人類評判中勝率達 63%——但效能指標看不出差異

Researchers compared two identical 1.2B-parameter AI models trained on the same data with one key difference: one used a new training method inspired by predictive coding (with precision-weighted gains and layer-specific gradient scaling), while the other used standard training. Here's the wild part—both models ended up with virtually identical performance scores, yet when 10 judges (7 humans + 3 AI systems) compared them blind, they preferred the new method 63.4% of the time. This suggests the training tweak creates something meaningful that standard metrics completely miss. With statistical significance at p = 1.98 × 10⁻⁵, this isn't a fluke. The implication? We might be measuring AI quality all wrong.