latentbrief
← Back to editorials

Editorial · General AI News

AI Benchmarks Have Reached Their Ceiling - And It’s a Problem Nobody Is Admitting

2h ago2 min brief

The AI industry has long celebrated benchmark after benchmark as proof of progress. But the latest round of metrics reveal a worrying truth: the models are hitting a wall. While performance in specific tasks like report drafting and policy creation has improved, the gains are diminishing - and the gap between what’s being promised and what’s actually delivered is growing.

The EQS AI Benchmark Volume 2, released earlier this year, shows that the top AI models now cluster closely together, with minimal differences in their compliance task performance. OpenAI's GPT-5.4 leads at 87.6%, followed by Google’s Gemini 3.1 Pro and Anthropic’s Claude Opus. The improvements are significant but not transformative - especially when compared to the hype surrounding these systems. The real issue is that while models are getting better, they’re not improving fast enough to justify the industry’s claims of revolutionary change.

This plateau in performance is happening at a time when the stakes are higher than ever. Compliance teams are increasingly relying on AI to handle multi-step workflows - from risk assessment to mitigation strategies. But as EQS Group’s Moritz Homann noted, the question isn’t whether AI can support these processes anymore. It’s how we design the systems around them. The human oversight and contextual understanding that should accompany these tools are often missing in discussions about model capabilities.

The problem lies in how benchmarks are designed. They focus on quantifiable metrics like accuracy and latency, ignoring the broader impact on human agency and critical thinking. This narrow approach lets the industry pretend that AI is a neutral tool rather than a system that can erode our ability to make decisions independently.

A new framework for evaluation is needed - one that measures not just what AI can do, but what it means for the people using it. Metrics like harm reduction, mental health outcomes, and long-term skill development should take center stage. Until then, any claims of AI reaching its full potential are nothing more than empty promises. The models may have reached their ceiling, but the real challenge is getting humanity to admit - let alone address - how far we’ve fallen behind.

Editorial perspective - synthesised analysis, not factual reporting.

Terms in this editorial

EQS AI Benchmark Volume 2
A benchmark that evaluates the performance of AI models in compliance tasks. It highlights how closely top models cluster together and their minimal differences in task performance, indicating a plateau in model improvements.
GPT-5.4
A version of OpenAI's GPT model that achieved 87.6% performance in the EQS AI Benchmark Volume 2, leading among other models like Google’s Gemini and Anthropic’s Claude.

If you liked this

More editorials.