A new benchmark called DiscoBench measured the performance of AI search agents on ambiguous queries. The result is surprising: the main issue isn't the ability to search for information, but the lack of clarifying questions asked to the user. In fact, models that search repeatedly without asking achieve only 51.9% accuracy. Even the best model stops at 43% overall.
However, when ambiguity is removed from queries, accuracy increases by up to 40 percentage points. This data is crucial. Therefore, the issue isn't strictly technological: it's conversational. An AI agent that doesn't know when to stop and ask produces worse answers than one that simply guesses.
At SHM Studio we closely monitor these dynamics because they directly impact the effectiveness of AI tools integrated into marketing and SEO workflows. Consequently, understanding the real limits of AI search agents is now a priority for any digital manager who wants to adopt these technologies consciously and measurably.
The benchmark that changes the perspective on AI search agents
For months, the debate on AI search agents has focused on the quality of the retrieved sources. However, recent research shifts the focus elsewhere. The benchmark DiscoBench , presented in July 2026, analyzes a different problem: what happens when a query is ambiguous and the agent has to decide whether to search further or ask the user for clarification.
The results, reported by The Decoder , are stark. Models that iterate searches without asking for clarification achieve only 51.9% accuracy. Paradoxically, they perform worse than those that simply guess the answer. Furthermore, the best-performing model doesn't exceed 43% overall accuracy.
So, the limit of AI search agents is not their ability to retrieve data. It's their ability to recognize when a question is too vague to produce a reliable answer.
Problem architecture: how a search agent works under pressure
An AI search agent, in its most common form, operates in cycles. It receives a query, breaks it down, performs multiple searches, and synthesizes the results. This process works well for precise queries. Conversely, for ambiguous queries, the system enters a counterproductive loop.
In particular, the problem emerges in three typical scenarios:
- Queries with multiple referents : the same word indicates different entities depending on the context (e.g., "Apple" as a company or as a fruit in a food context).
- Queries with implicit intent : the user assumes a context that the agent doesn't know.
- Queries with undefined scope : the request is too broad, and the agent doesn't know where to narrow down the search.
In these cases, searching more doesn't help. On the contrary, it worsens the result. Therefore, the critical variable becomes the agent's ability to recognize ambiguity and break the cycle to ask.
According to recent research by McKinsey on the adoption of AI in companies, the quality of AI tools' output increasingly depends on the quality of human-machine interaction. This finding aligns perfectly with what emerges from DiscoBench.
The 40-point leap: what happens without ambiguity
DiscoBench also includes a control condition. When queries are reformulated to eliminate ambiguity, the agents' accuracy increases by up to 40 percentage points. This is the most relevant finding of the entire study.
This means the underlying technology works. The search engine, the language model, the summarization capability: everything is operational. The problem lies upstream. It concerns the intent comprehension phase, not the execution phase.
Consequently, those working with AI search agents in professional contexts should focus on two areas. First, the quality of the initial prompt and the structure of the queries. Second, configuring the agent so it knows when to ask, not just how to search.
This has direct implications for those using AI tools in their workflows. SEO and Digital marketing .
SME use cases: where this limitation is really felt
For an SME or mid-market company, AI search agents are often used for competitive research, market analysis, or content production support. In these contexts, ambiguity is the norm, not the exception.
For example, a query like “competitor analysis in the sustainable packaging sector” can refer to direct, indirect, geographically limited, or global competitors. An agent that doesn’t ask what scope to apply produces an output that mixes incompatible analysis levels.
Similarly, in activities of SEO copywriting AI-assisted, an ambiguous request about tone or target audience generates content that misses the mark. Despite this, many teams continue to use these tools without an intent validation procedure.
Moreover, the problem is amplified when the agent is integrated into automated workflows. In that case, there is no human operator ready to correct the course. The incorrect output propagates throughout the process.
Trade-offs: asking too much or not asking enough
There is a delicate balance to find. An agent that asks for clarification on every query becomes unusable. Conversely, one that never asks produces unreliable results on complex queries.
Research on DiscoBench doesn't yet indicate the optimal intervention threshold. However, it suggests that current models systematically err towards the latter extreme: they try to ask instead. This behavior is likely a result of training, aimed at maximizing immediate responses.
According to an analysis by Harvard Business Review on the operational use of AI, the tendency of models to
Related articles
Discover more articles exploring similar topics, selected to offer you a more complete and stimulating perspective. Each piece of content is carefully chosen to enrich your experience.