Summary
Recent AI model benchmarks, such as for Alibaba's Qwen 3.8-Max and Claude Opus 5, highlight that raw scores can be misleading due to varying time and token budgets. The article argues for using "cost per successful task" as a more accurate metric, emphasizing that token prices alone no longer predict the actual bill for reasoning models.