AI search agents and ambiguity: the problem is asking
A new benchmark called DiscoBench measured the performance of AI search agents on ambiguous queries. The result is surprising: the main problem isn’t the ability to search for information, but the failure to ask the user for clarification. In fact, models that search repeatedly without asking for clarification achieve an accuracy of only 51.9%. Even the best model only reaches an overall accuracy of 43%.
However, when ambiguity is removed from queries, accuracy increases by up to 40 percentage points. This data is crucial. Therefore, the bottleneck is not strictly technological; it is conversational. An AI agent that doesn't know when to stop and ask produces worse answers than one that simply guesses.
In SHM Studio, we monitor these dynamics carefully, as they directly impact the effectiveness of AI tools integrated into marketing and SEO workflows. Consequently, understanding the real limitations of AI search agents is a priority today for any digital manager who wants to adopt these technologies in a conscious and measurable way.
The benchmark that changes the perspective on AI search agents
For months, the debate surrounding AI search agents has focused on the quality of retrieved sources. However, recent research shifts the focus elsewhere. The benchmark DiscoBench, launched in July 2026, analyzes a different problem: what happens when a query is ambiguous and the agent has to decide whether to search further or ask the user for clarification.
The results, reported by The Decoder, are clear. Models that iterate through searches without asking for clarification achieve an accuracy of only 51.9%. Paradoxically, they perform worse than those that simply guess the answer. Furthermore, the best-performing model does not exceed 43% in overall accuracy.
So, the limit of AI search agents isn't their ability to retrieve data. It's their ability to recognize when a question is too vague to produce a reliable answer.
Problem Architecture: How a Search Agent Works Under Pressure
An AI search agent, in its most common form, operates in cycles. It receives a query, breaks it down, performs multiple searches, and synthesizes the results. This process works well for precise queries. Conversely, for ambiguous queries, the system enters a counterproductive loop.
In particular, the problem arises in three typical scenarios:
- Query with multiple referencesthe same word indicates different entities depending on the context (e.g., “Apple” as a company or as a fruit in a food context).
- Query with implicit intentThe user assumes a context that the agent does not know.
- Query with undefined scopeThe request is too broad, and the agent doesn't know where to limit the search.
In these cases, searching more doesn't help. In fact, it worsens the result. Therefore, the critical variable becomes the agent's ability to recognize ambiguity and break the cycle to ask.
According to recent research from McKinsey Regarding the adoption of AI in business, the quality of AI tools' output increasingly depends on the quality of human-machine interaction. This finding aligns perfectly with what emerges from DiscoBench.
The 40-point jump: what happens unambiguously
DiscoBench also includes a control condition. When queries are reformulated to eliminate ambiguity, agent accuracy increases by up to 40 percentage points. This is the most relevant finding of the entire study.
It means the underlying technology works. The search engine, the language model, the synthesis capability: everything is operational. The problem is upstream. It concerns the intent comprehension phase, not the execution phase.
Consequently, those working with AI search agents in professional contexts should focus on two areas. First, the quality of the initial prompt and query structure. Second, configuring the agent so it knows when to ask, not just how to search.
This has direct implications for those using AI tools in their workflows SEO e digital marketing.
SME Use Cases: Where This Limit Really Comes Into Play
For an SME or mid-market company, AI search agents are often used for competitive research, market analysis, or content production support. In these contexts, ambiguity is the norm, not the exception.
For example, a query like “competitor analysis in the sustainable packaging sector” can refer to direct, indirect, geographically limited, or global competitors. An agent that doesn't ask for the scope to apply will produce output that mixes incompatible levels of analysis.
Analogously, in activities of SEO copywriting Assisted by AI, an ambiguous request regarding tone or target audience generates content that misses the mark. Despite this, many teams continue to use these tools without an intent validation process.
In addition to this, the problem is amplified when the agent is integrated into automated workflows. In that case, there is no human operator ready to correct it. The incorrect output propagates throughout the entire process.
Trade-off: Asking for too much versus not asking enough
There's a delicate balance to strike. An agent that asks for clarification on every query becomes unusable. Conversely, one that never asks produces unreliable results on complex queries.
Research on DiscoBench does not yet indicate the optimal intervention threshold. However, it suggests that current models systematically err toward the second extreme in their response: they seek to ask instead. This behavior is likely a result of training, which is oriented toward maximizing immediate response.
According to an analysis by Harvard Business Review on the operational use of AI, the tendency of models to
News Categories
Related articles
Discover other articles that explore similar topics in depth, selected to give you a more complete and stimulating view. Each piece of content is carefully chosen to enrich your experience.