TL;DR
Hex introduces DataBench, a 100-task benchmark for evaluating AI agents on real-world analytics work, revealing models struggle with judgment and avoid admitting uncertainty.
Key Points
- GPT-5.6 Luna dominates cost-efficiency, achieving near-Sol performance at 1/14th the cost on Pareto frontier
- Claude Fable 5 only model avoiding performance regression at maximum effort levels; Opus 5 talks itself into wrong answers
- Models score 75% on evidence-gathering Q&A tasks but only 54% on trap tasks requiring deeper reasoning and judgment
- Frontier models rarely admit uncertainty; only Opus 5 passes collections-call-list trap requiring skepticism of plausible-but-wrong evidence
Why It Matters
Existing analytics benchmarks (Spider 2.0, DABstep) report 85-90% accuracy but test toy problems, not real work. DataBench exposes critical gaps: frontier models manufacture false confidence, struggle with ambiguous prompts, and regress when given more compute. This matters for anyone deploying AI agents on actual business analytics—you need skepticism and human oversight, not blind trust in high benchmark scores.
Source: hex.tech