UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do

In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budget. On software engineering tasks, success rates jumped about 25 percent when the token budget was increased tenfold. Newer models benefit the most. Depending on the token budget, actual progress at the frontier is about…
This is a summary curated by AIFuture. Read the complete article at the original source:
Read the full story on The Decoder