Pull down to go back
MiniMax2.7 Underperforms on Terminal Bench—Local Testing Shows It's Actually Worse Than M2.5

MiniMax2.7 Underperforms on Terminal Bench—Local Testing Shows It's Actually Worse Than M2.5

MiniMax2.7 在 Terminal Bench 表現慘淡——實測竟然比 M2.5 還差

Someone just ran a full Terminal-Bench 2.0 test with MiniMax-M2.7 locally on a Mac Studio M3 Ultra, and the results are disappointing. After 445 trials, the model hit only 41.3% accuracy—which is actually *worse* than the 42.7% they got with the older M2.5 on the exact same hardware. Out of 434 trials, 184 solved and 250 failed, with 187 timeout errors (the model just ran out of time, not crashed). The speed was also sluggish at 10-17 tokens per second. If you're thinking about using MiniMax2.7 for agent coding tasks in Claude, this benchmark suggests you might want to stick with M2.5 or look elsewhere—the newer version doesn't seem to be delivering the upgrade people expected.