A new benchmark called DiscoBench tested AI search agents on ambiguous queries. The result is surprising: the problem isn't the ability to search for information, but the inability to ask the user for clarification when the question is poorly defined. Therefore, models that repeat searches instead of doing follow-ups achieve just 51.9% accuracy — worse than guessing directly.
Plus, the best model tested stops at 43% overall accuracy. However, when ambiguity is removed from the queries, accuracy jumps up by 40 percentage points. This data radically changes the perspective on how to evaluate and integrate AI agents into business workflows. In particular, anyone adopting AI-based solutions for research, content, or customer intelligence needs to rethink query design and workflows.
We at SHM Studio we closely follow the evolution of AI agents applied to digital marketing. Consequently, we believe this data has practical implications for Italian marketing managers who are evaluating or already using AI tools in their processes. In this analysis, we dive deep into the problem's architecture, practical use cases, and operational guidelines.
The benchmark that changes the narrative on AI agents
For months, the debate on AI search agents has focused on the quality of the retrieved sources. However, recent research shifts the focus elsewhere. The benchmark DiscoBench , analyzed by The Decoder , shows that the real bottleneck is not information retrieval. It is managing ambiguity in the initial query.
Basically, when a user makes a vague or polysemous request, the AI agent should ask for clarification. Instead, most models tend to iterate searches, hoping to converge on an answer. This approach turns out to be counterproductive: models that repeat searches get 51.9% accuracy. On the other hand, those that simply guess do better.
So, the problem is not computational. It is communicative. And this distinction has direct consequences for those who integrate AI agents into marketing and business intelligence processes.
Problem architecture: why the agent doesn't ask
To understand the phenomenon, it's helpful to look at how a typical AI search agent works. The model receives a query, generates a search plan, executes calls to engines or databases, and synthesizes the results. This cycle is optimized for retrieval completeness , not for the semantic precision of the request.
Therefore, when the query is ambiguous — for example
Related articles
Discover more articles exploring similar topics, selected to offer you a more complete and stimulating perspective. Each piece of content is carefully chosen to enrich your experience.