WORLDTECH NEWS Global technology intelligence.Contact
← Back to WORLDTECH
AI SINGLE SOURCE

AI agent teams waste massive tokens for barely measurable quality gains, research finds

Detailed close-up view of a smartphone with a textured back cover emphasizing modern design and technology.
Illustrative photo.Photo by Pixabay on Pexels

What happened

Plus AI in practice Copy the url to clipboard Share this article Go to comment section AI agent (AI that carries out multi-step tasks rather than answering one question) teams waste massive tokens (the small pieces of text a model reads and writes) for barely measurable quality gains, research finds Matthias Bastian View the LinkedIn Profile of Matthias Bastian Oct 11, 2026 AI agent teams deliver almost no better results than single agents, research finds. Evals company Vals AI tested GPT-6 (large language model by OpenAI) Sol and Claude Opus 5.5 on the "Vibe Code Bench," both solo and as teams, at two reasoning levels: medium and maximum reasoning effort.

The results suggest that the extra cost of agent teams isn't worth it in most cases, especially when models are already running at full compute. Teams of AI agents barely outperform solo agents but cost up to 5.1x more, according to Vals AI.

Only one out of four tests with GPT-6 Sol and Claude Opus 5.5 showed a measurable gain. Anthropic's own data backs this up: beyond ten agents, quality plateaus while token costs keep climbing.

The teams cost between 1.8x and 5.1x more than single agents. Out of four comparisons between teams and solo agents, only one showed a statistically significant improvement: GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher. At maximum reasoning, the team setup gave neither Sol nor Opus 5.5 any real advantage.

Key facts

  • Evals company Vals AI — tested: GPT-6 Sol and Claude Opus 5.5 on the "Vibe Code Bench," both solo and as teams, at two reasoning levels: medium and maximum reasoning effort

Sources & evidence