Washington | 19°C (overcast clouds)
Frontier AI Stumbles on Real‑World Workplace Tests

New Berkeley study finds cutting‑edge agents can’t handle most job tasks

A UC‑Berkeley evaluation of the latest AI models shows they flunk a massive exam of real‑world work, with the best system passing only about a quarter of tasks.

Artificial‑intelligence spending has already topped $1.6 trillion, and the hype machines keep promising that the next wave will finally make AI useful at work. The reality, at least for now, looks a lot messier.

Researchers at the University of California, Berkeley’s Center for Responsible, Decentralized Intelligence put together what they call the “Agents’ Last Exam” (ALE). Think of it as a massive, 1,500‑item quiz covering 55 different occupations – from software engineering and graphic design to maritime engineering, agriculture, and public‑health operations. The idea was simple: see whether the most advanced AI agents can actually do the jobs we’re already handing them.

They tested a line‑up of “closed” models – proprietary systems that most people don’t get to poke at directly – including Anthropic’s Fable 5, OpenAI’s brand‑new GPT‑5.5, Cursor’s Composer 2.5, and Google’s Gemini 3.1 Pro. For completeness they also threw in two open‑source models from Chinese developers.

The results were sobering. Every model flunked the exam in its own way. GPT‑5.5, which topped the pack, managed to pass just 24 percent of the tasks overall. That sounds better than zero, but it still means three‑quarters of the problems were unsolvable for a system that’s supposed to be the cutting edge.

When the researchers dug into the harder tier of questions – those demanding sustained reasoning, deep domain expertise, and long‑term execution – the picture got even bleaker. No agent, not even Fable 5, achieved a single correct answer. In other words, the toughest real‑world problems remain completely out of reach for today’s frontier AI.

Cost is another angle the study explored. Fable 5, for example, delivered performance roughly on par with GPT‑5.5 and Composer 2.5, but it cost somewhere between four to twelve times more per completed task. In a world where businesses chase efficiency, that price tag is hard to swallow.

Still, the authors caution against writing off AI’s impact entirely. “Even if current pass rates remain relatively low, occupations dominated by routine and well‑defined procedures are likely to experience disruption first, while decision‑intensive roles will remain more resilient for longer,” said Dawn Song, a co‑author of the paper.

In plain English: the jobs that are highly repetitive – data entry, basic report generation, simple code linting – could feel the squeeze soon, even if the AI behind them isn’t perfect. Roles that require nuanced judgment, creativity, or deep expertise are likely to stay safer for the foreseeable future.

The takeaway? The AI revolution isn’t a sudden, flawless leap. It’s a staggered, uneven climb, and for now the most advanced agents are still figuring out how to get up the first few rungs without tripping.

Comments 0
Please login to post a comment. Login
No approved comments yet.

Editorial note: Nishadil may use AI assistance for news drafting and formatting. Readers can report issues from this page, and material corrections are reviewed under our editorial standards.