- The core problem: measuring AI like traditional software
- Framework Architecture: Three Levels of Reading
- How to calculate useful work per dollar in practice
- High-value use cases for B2B and retail marketing
- Efficiency vs. Scaling: When to Accelerate and When to Stop
- The construction site still open: governance and human control
- Operational metrics to monitor over time
- Mistakes to Avoid During Budget Allocation
- Prospects for 2027: Towards Structural AI Budgets
The era of agentive AI changes the game for those managing tech budgets. It's no longer enough to count software licenses or saved hours. We need to measure the useful work produced for every euro invested. This is the central principle that OpenAI has articulated in its recent framework for businesses.
However, for Italian SMEs, the conceptual leap is even more challenging. Many companies find themselves evaluating AI tools without adequate metrics. Consequently, investments risk being understated or, conversely, scattered across low-impact use cases. In particular, the challenge is not technological; it is strategic and organizational.
We of SHM Studio We work daily with marketing managers and digital managers of Italian companies who are exactly in this transition phase. Therefore, in this article, we propose an operational reading of the OpenAI framework, adapted to the reality of SMEs and the B2B mid-market. Finally, we indicate the concrete metrics to monitor and the most common errors to avoid during the scaling phase.
The fundamental problem: measuring AI like traditional software
For years, companies have evaluated software investments in terms of licensing costs, implementation hours, and reduction of manual labor. This approach worked in a context where digital tools performed predefined tasks. However, agentive AI introduces a substantial discontinuity.
An AI agent does not perform a task: orchestra of autonomous decision sequences to achieve a goal. As a result, the value generated is not linear with respect to the cost. It can be much higher—or zero, if the use case is poorly selected.
OpenAI has addressed this issue directly in its Framework for AI Investment Management in the Agent Era, published in July 2026. The document introduces the concept of value for money as a guiding metric. Therefore, the question is not «how much does AI cost?» but «how much useful work does each euro spent produce?»
Framework Architecture: Three Levels of Reading
The framework is structured into three distinct levels. First, high-value workflows must be identified. Next, the operational efficiency of agents on those workflows is measured. Finally, only what has demonstrated measurable impact is scaled.
This approach avoids the most common trap: implementing AI on marginal processes because they are «easier,» accumulating costs without generating strategic value. In fact, according to research from McKinsey on the Value of AI in Business, the companies that achieve the best results focus investments on a limited number of high-impact use cases, rather than distributing them across many simultaneous pilot projects.
Furthermore, the framework distinguishes between process efficiency (doing the same things faster) and Outcome effectiveness (achieving results that were not previously possible). SMEs tend to measure only the first dimension. However, it is often the second that justifies the most significant investments.
How to calculate useful work per dollar in practice
The metric value for money requires an operational definition of «useful work.» This varies for each organization and each business function. For example, for a marketing team, useful work could be the number of qualified content pieces produced, leads generated, or campaigns autonomously optimized by the agent.
To reliably calculate this metric, three elements are needed. First, a pre-AI baseline of the cost per unit of output. Second, a clear definition of «acceptable quality» for agentive output. Third, a continuous monitoring system that detects quality degradation over time.
We of SHM Studio We often observe that Italian SMEs invest in the initial implementation but neglect the third element. As a result, the initial ROI is not maintained over time. Therefore, continuous monitoring is not an optional activity: it is an integral part of the investment.
High-Value Use Cases for B2B and Retail Marketing
Which workflows deserve priority in AI budget allocation? The answer depends on the industry and organizational structure. However, some categories regularly emerge in the Italian B2B and retail contexts.
In B2B marketing, the most productive workflows involve automated lead qualification, scalable content personalization, and real-time optimization of paid campaigns. For example, an AI agent integrated with the LinkedIn campaign can analyze engagement signals and reallocate budget to the most responsive segments without manual intervention.
In retailing, high-value use cases include dynamic promotion management, personalized transactional emails, and inventory planning support. Additionally, integration with systems advertising on Google Ads allows for optimizations that a human team could not perform at the same speed and granularity.
In both contexts, the principle remains the same: Select workflows based on output value, not ease of implementation. This requires a strategic discussion even before choosing the technological tools.
Efficiency vs. Scaling: When to Accelerate and When to Stop
The OpenAI framework introduces an important distinction between the efficiency phase and the scaling phase. Many organizations burn through budgets because they skip the first and go straight to the second.
The efficiency phase serves to validate that an AI agent actually produces useful work on a specific use case. This phase should be brief, contained, and measurable. Conversely, scaling assumes that validation has already occurred and that the organization has the operational capacity to handle larger volumes.
According to Gartner on Generative AI, one of the main causes of failure in enterprise AI projects is precisely the absence of this distinction. Companies scale before they have optimized, multiplying costs without multiplying value. Therefore, process discipline is more decisive than technological choice.
The construction site still open: governance and human control
The autonomy of AI agents raises a governance issue that the framework addresses pragmatically. Agents cannot be left to operate unsupervised, especially in the initial stages. However, excessive supervision negates the benefits of efficiency.
The break-even point depends on the level of risk associated with the workflow. For example, an agent that generates content drafts requires light supervision. Conversely, an agent that manages advertising budgets or interacts with clients in real-time requires more structured controls.
For SMEs, this translates into a concrete operational recommendation: define Graduated levels of autonomy for each agentive workflow, with human review checkpoints proportional to the risk. This approach is compatible with the limited resources typical of medium-sized organizations. Furthermore, it facilitates compliance with European AI regulations, which require traceability of automated decisions.
The activities of AI-assisted copywriting, for example, typically fall into the low-risk category. Therefore, they can be scaled with reduced supervision once the output quality has been validated.
Operational metrics to monitor over time
Defining the right metrics is a prerequisite for effective AI budget management. Below are the main categories of indicators to include in a monitoring dashboard.
- Useful work rate percentage of generative outputs that meet defined quality criteria, without significant human review.
- Cost per qualified output Total cost (licenses + supervision + infrastructure) divided by the number of qualified outputs produced in the period.
- Time to value time elapsed between the activation of an agent workflow and the first measurable useful output.
- Escalation rate frequency with which the agent transfers control to a human operator. A high rate signals a use case that is not yet mature for automation.
- Quality drift variation in the quality of outputs over time. Agents tend to degrade if not updated or if the context changes.
These indicators should be integrated into the processes of Digital Marketing Management and reviewed monthly. Additionally, it is useful to compare them with industry benchmarks when available.
Mistakes to Avoid During Budget Allocation
On-the-ground experience highlights some recurring errors that erode the ROI of AI investments. Specifically, these errors are concentrated in the planning phase and the first few weeks of implementation.
The first error is underestimate integration costs. The cost of the AI license is often just a fraction of the total cost. Integration with existing systems, team training, and ongoing monitoring account for the bulk of the actual investment.
The second mistake is select use cases based on technological availability rather than its strategic value. Many vendors offer preconfigured use cases. However, these may not align with the organization's specific priorities.
The third mistake is do not define a baseline before implementation. Without pre-AI data, it is impossible to measure the improvement. As a result, ROI remains a qualitative claim rather than a verifiable figure.
Finally, the fourth error is Scale too soon., as already highlighted. The pressure to demonstrate rapid results often pushes teams to expand the use of agents before the validation phase is completed. Therefore, it is useful to set explicit progression criteria before starting the project.
To further explore how to structure an AI strategy suitable for Italian SMEs, the team SHM Studio is available for an initial consultation. Furthermore, our blog gathers updated analyses on the sector's main developments.
Prospects for 2027: Towards Structural AI Budgets
Projections for the 2027–2028 period indicate that the portion of the marketing budget allocated to agent-based AI tools will grow significantly. However, this growth will not be uniform. Organizations that have established robust governance and reliable metrics by 2026 will be able to scale efficiently. Conversely, those that have accumulated unvalidated implementations will find themselves having to streamline their portfolio before they can move forward.
Therefore, 2026 is the year to lay the methodological groundwork. It is not yet time for widespread scaling; it is time for discipline. SMEs that invest today in defining the right metrics and governing agent workflows are positioning themselves to reap the benefits in the coming quarters.
For those who wish to delve deeper into the operational implications for their organization, the team SHM Studio is reachable here. We offer a free initial assessment to identify high-value workflows and estimate the potential ROI of generative AI within your company's specific context.
Related articles
Discover other articles that explore similar topics in depth, selected to give you a more complete and stimulating view. Each piece of content is carefully chosen to enrich your experience.