- The breaking point: what Sutton really said
- Problem architecture: generation without verification
- AlphaGo and AlphaProof: the model with feedback loop
- SME Use Cases: where this distinction is already relevant
- The still-open construction site: where research has not responded
- Operational trade-offs: generative vs. evaluative in business decisions
- Recommended decision: how to navigate today
Richard Sutton, winner of the 2024 Turing Award, has raised a central question in the AI debate. His thesis is clear: pure generative AI systems are incapable of doing real science. The reason is structural. Without an internal mechanism to evaluate their own results, these systems cannot distinguish a real discovery from a plausible hallucination.
Furthermore, Sutton points to an alternative, already working model. Systems like AlphaGo and AlphaProof integrate autonomous evaluation loops. As a result, they can be genuinely creative and produce verifiable outputs. In contrast, a generative model that produces text or code without self-evaluation generates ephemeral novelty: interesting on the surface, but scientifically unreliable. Therefore, the distinction between generative AI and AI with feedback loops is not just technical — it's epistemological.
We at SHM Studio we are carefully following this debate. In particular, the implications for Italian SMEs that are considering investments in tools of Artificial intelligence are concrete and immediate. Understanding what generative AI can and cannot do is the first step in choosing the right solutions.
The breaking point: what Sutton really said
Richard Sutton is one of the most authoritative voices in the field of artificial intelligence. Winner of the Turing Award, father of reinforcement learning modern, he recently expressed a position that deserves in-depth analysis. According to Sutton, pure generative AI is not capable of doing real science . The reason is structural, not contingent.
The central problem is the lack of self-assessment. A generative system produces output — text, code, images, hypotheses — but it doesn't have an internal mechanism to verify its correctness. Therefore, the novelty it generates is, in Sutton's words, 'ephemeral': it appears, seems plausible, then vanishes without leaving a verifiable trace.
Therefore, the distinction Sutton proposes is not about computational power. It is about the system’s epistemic architecture. This perspective has direct implications for anyone evaluating the adoption of tools based on Artificial intelligence in professional settings.
Problem architecture: generation without verification
To understand Sutton's critique, it's helpful to analyze how a pure generative system works. Models like GPT-4, Claude, or Gemini operate on a statistical principle: given a context, they predict the most likely token. They then produce coherent, fluent, often convincing sequences.
However, this consistency is syntactic and probabilistic, not semantic or truth-based. The model doesn't know if what it wrote is true. It doesn't have an internal oracle that compares the output with reality. Consequently, it can produce false statements with the same fluency as correct ones.
In the scientific field, this is a critical limit. Scientific discovery requires a cycle: hypothesis, testing, evaluation, revision. Without the step of autonomous evaluation, the cycle breaks. In fact, a system that cannot falsify its own hypotheses cannot do science in the Popperian sense of the term. This point is also explored in a recent analysis by MIT Technology Review on the structural limits of language models .
AlphaGo and AlphaProof: the model with feedback loop
Sutton points to an alternative direction that is already feasible. Systems like DeepMind's AlphaGo and AlphaProof integrate an internal evaluation mechanism. In AlphaGo, each move is evaluated by a value function that estimates the probability of winning. Therefore, the system does not just generate plausible moves: it compares them, ranks them, and discards them.
AlphaProof, developed for Olympic mathematics, works analogously. It generates proofs and formally verifies them. Consequently, it can distinguish a correct proof from an incorrect one. This evaluation loop is what Sutton calls the prerequisite for true computational creativity.
Furthermore, this approach is not new. It is the logic of the reinforcement learning : an agent acts, receives a reward signal, updates its policy. The reward signal is, in essence, the evaluation. Without it, the agent does not learn — it merely generates. As reported by DeepMind on its research portal , systems with structured feedback show qualitatively different reasoning capabilities.
SME Use Cases: where this distinction is already relevant
For an Italian SME, the distinction between pure generative AI and AI with an evaluation loop is not just academic. There are operational contexts where this difference translates into concrete risks.
First, let's consider AI-assisted copywriting . A generative model produces fluent text, but it can insert incorrect data, non-existent citations, or unverified statements. Therefore, without a human review process or an automatic verification system, the risk of publishing inaccurate content is high.
Similarly, in market analysis or report generation, a system without self-evaluation can produce seemingly solid insights that lack empirical grounding. Consequently, strategic decisions based on these outputs can be skewed. This also applies to applications in the field of Digital marketing , where data quality is crucial.
Conversely, systems that integrate verification — for example through Retrieval-Augmented Generation (RAG) with certified sources, or pipelines with automatic validation — significantly reduce this risk. Therefore, the choice of the right tool depends on understanding this architecture.
The still-open construction site: where research has not responded
Sutton's position is stimulating, but not without gray areas. Some researchers object that the latest generation language models show emergent capabilities of self-consistency checking : they compare alternative answers and select the most coherent one. However, this is not equivalent to a true epistemic evaluation.
Furthermore, the line between "generation" and "evaluation" is not always clear in engineering practice. Systems like OpenAI's o1 incorporate chains of internal reasoning (" chain-of-thought ) that simulate a verification process. Despite this, Sutton would argue that without an external reality signal — an environment, a formal oracle, an empirical test — this too remains sophisticated probabilistic processing, not genuine evaluation.
Therefore, the debate is open. The scientific community has not yet reached a consensus on where to draw the line. As highlighted by a recent analysis on Harvard Business Review on the strategic use of generative AI , the distinction between a support tool and an autonomous agent is central to any business adoption decision.
Operational trade-offs: generative vs. evaluative in business decisions
From a practical standpoint, the two architectures have very different cost and applicability profiles. Pure generative systems are accessible, inexpensive, and quick to integrate. Therefore, they are suitable for content production tasks, brainstorming, document summarization, and support for LinkedIn campaign management or to writing creative briefs.
Systems with evaluation loops are more complex to build and maintain. They require a testing environment, a well-defined reward signal, and often structured proprietary data. As a result, they are better suited for contexts where correctness is critical: diagnostics, research, optimization of google ads campaigns with automatic feedback, or recommendation systems with verifiable performance metrics.
Finally, there is an increasingly viable middle ground: integrating generative systems with external validation layers. This is the direction in which major enterprise platforms are moving. We at SHM Studio we are observing this evolution with attention, particularly for the implications on services of SEO and web development AI-assisted.
Recommended decision: how to navigate today
Sutton’s thesis offers a practical criterion for evaluating any AI tool. First and foremost, one must ask: is this system capable of evaluating its own outputs? Does it have a feedback mechanism that goes beyond internal consistency?
For Italian SMEs operating in sectors where precision is critical — manufacturing, consulting, healthcare, finance — the answer to this question should guide the choice of tool. In particular, it is advisable to favor solutions that integrate automatic validation or that provide for a structured human review process.
For creative, marketing, or communication tasks, pure generative systems remain effective and convenient tools. However, even in these contexts, human supervision isn't optional — it's an integral part of the process. To delve deeper into how to correctly integrate these tools into your digital strategy, you can contact the SHM Studio team or explore the resources available in the Blog .
Related articles
Discover more articles exploring similar topics, selected to offer you a more complete and stimulating perspective. Each piece of content is carefully chosen to enrich your experience.