- The core problem: measuring AI like traditional software was measured
- Framework architecture: three levels of reading
- Useful work per dollar: how to calculate it in practice
- High-value use cases for B2B and retail marketing
- Efficiency vs. scaling: when to speed up and when to pause
- The ongoing construction site: governance and human control
- Operational metrics to track over time
- Mistakes to avoid during the budget allocation phase
- Outlook for 2027: toward structural AI budgets
The era of agentive AI changes the game for those managing technology budgets. Counting software licenses or saved hours is no longer enough. It's necessary to measure the useful work produced for every euro invested . This is the core principle that OpenAI articulated in its recent enterprise framework.
However, for Italian SMEs, the conceptual leap is even more challenging. Many companies find themselves evaluating AI tools without adequate metrics. Consequently, investments risk being undersized — or, conversely, scattered across low-impact use cases. In particular, the challenge is not technological: it is strategic and organizational.
We at SHM Studio We work daily with marketing managers and digital leaders from Italian companies who are at this exact transition point. Therefore, in this article, we propose an operational reading of the OpenAI framework, adapted to the reality of SMEs and the B2B mid-market. Finally, we indicate the concrete metrics to monitor and the most common mistakes to avoid during the scaling phase.
The core problem: measuring AI like traditional software was measured
For years, companies have evaluated software investments in terms of license costs, implementation hours, and reduction of manual labor. This approach worked in a context where digital tools performed predefined tasks. However, agentive AI introduces a substantial discontinuity.
An AI agent doesn't just run a task: orchestrates sequences of autonomous decisions to achieve a goal. Consequently, the value generated is not linear relative to the cost. It can be much higher — or zero, if the use case is poorly selected.
OpenAI tackled this topic directly in its framework for managing AI investments in the agentic era , published in July 2026. The document introduces the concept of useful work per dollar as the guiding metric. Therefore, the question is not «how much does AI cost?» but «how much useful work does every euro spent produce?»
Framework architecture: three levels of reading
The framework is divided into three distinct levels. First of all, high-value workflows need to be identified. Next, the operational efficiency of the agents on those workflows is measured. Finally, only what has proven a measurable impact is scaled.
This approach avoids the most common trap: implementing AI on marginal processes because they are "easier," accumulating costs without generating strategic value. In fact, according to research by McKinsey on the value of AI in business , companies that achieve the best results focus their investments on a limited number of high-impact use cases, rather than spreading them across many simultaneous pilot projects.
Furthermore, the framework distinguishes between process efficiency (doing the same things faster) and outcome effectiveness (achieving results that were previously impossible). SMEs tend to measure only the first dimension. However, it is often the second that justifies the most significant investments.
Useful work per dollar: how to calculate it in practice
The metric useful work per dollar requires an operational definition of "useful work." This varies for each organization and each business function. For example, for a marketing team, useful work could be the number of qualified content pieces produced, leads generated, or campaigns optimized autonomously by the agent.
To reliably calculate this metric, three elements are needed. First, a pre-AI baseline of the cost per unit of output. Second, a clear definition of 'acceptable quality' for agentive output. Third, a continuous monitoring system that detects quality degradations over time.
We at SHM Studio we often observe that Italian SMEs invest in the first implementation, but neglect the third element. Consequently, the initial ROI is not sustained over time. Therefore, continuous monitoring is not an optional activity: it is an integral part of the investment.
High-value use cases for B2B and retail marketing
Which workflows deserve priority in AI budget allocation? The answer depends on the sector and organizational structure. However, some categories emerge regularly in Italian B2B and retail contexts.
In B2B marketing , the most productive workflows involve automated lead qualification, at-scale content personalization, and real-time paid campaign optimization. For instance, an AI agent integrated with the LinkedIn campaigns it can analyze engagement signals and reallocate the budget toward the most responsive segments without manual intervention.
In Retail , high-value use cases include dynamic promotion management, transactional email personalization, and inventory planning support. Furthermore, integration with advertising on Google Ads allows for optimizations that a human team could not execute at the same speed and granularity.
In both contexts, the principle stays the same: select workflows based on output value, not ease of implementation . This requires a strategic conversation even before choosing the tech tools.
Efficiency vs. scaling: when to speed up and when to pause
The OpenAI framework introduces an important distinction between the efficiency phase and the scaling phase. Many organizations burn through budgets because they skip the first and jump straight to the second.
The efficiency phase serves to validate that an AI agent actually produces useful work on a specific use case. This phase should be short, contained, and measurable. Conversely, scaling assumes that validation has already occurred and that the organization has the operational capacity to handle larger volumes.
According to Gartner on agentic AI , one of the main causes of failure in enterprise AI projects is precisely the absence of this distinction. Companies scale before they've optimized, multiplying costs without multiplying value. Therefore, process discipline is more decisive than technology choice.
The ongoing construction site: governance and human control
The autonomy of AI agents raises a governance issue that the framework addresses pragmatically. Agents cannot be left to operate without supervision, especially in the early stages. However, excessive supervision nullifies efficiency benefits.
The break-even point depends on the risk level associated with the workflow. For example, an agent generating content drafts requires light supervision. Conversely, an agent managing advertising budgets or interacting with customers in real-time requires more structured controls.
For SMEs, this translates into a concrete operational recommendation: defining graduated autonomy levels for each agentive workflow, with human review checkpoints proportional to risk. This approach is compatible with the limited resources typical of medium-sized organizations. Furthermore, it facilitates compliance with European AI regulations, which require traceability of automated decisions.
The activities of AI-assisted copywriting , for example, typically fall into the low-risk category. Therefore, they can be scaled with reduced supervision once the output quality is validated.
Operational metrics to track over time
Defining the right metrics is a necessary condition for effective AI budget management. Below are the main categories of indicators to include in a monitoring dashboard.
- Useful work rate: percentage of agentic outputs that meet the defined quality criteria, without significant human review.
- Cost per qualified output: total cost (licenses + supervision + infrastructure) divided by the number of qualified outputs produced in the period.
- Time-to-value: time elapsed between the activation of an agentic workflow and the first measurable useful output.
- Escalation rate: frequency with which the agent transfers control to a human operator. A high rate signals a use case not yet mature for automation.
- Quality drift: variation in output quality over time. Agents tend to degrade if they are not updated or if the context changes.
These indicators should be integrated into the processes of digital marketing management and reviewed on a monthly basis. Furthermore, it is useful to compare them with industry benchmarks when available.
Mistakes to avoid during the budget allocation phase
Field experience highlights some recurring mistakes that erode the ROI of AI investments. In particular, these errors focus on the planning phase and the first few weeks of implementation.
The first mistake is underestimating integration costs . The cost of the AI license is often a fraction of the total cost. Integration with existing systems, team training, and continuous monitoring represent the largest part of the real investment.
The second mistake is choose use cases based on tech availability rather than strategic value. Many vendors offer pre-configured use cases. However, these may not align with the organization's specific priorities.
The third mistake is not defining a baseline before implementation. Without pre-AI data, it's impossible to measure improvement. Consequently, ROI remains a qualitative claim rather than verifiable data.
Finally, the fourth mistake is scaling too soon , as already highlighted. The pressure to demonstrate rapid results often pushes teams to expand agent use before the validation phase is complete. Therefore, it is helpful to set explicit advancement criteria before starting the project.
To learn more about how to structure an AI strategy tailored for Italian SMEs, the team at SHM Studio is available for an initial consultation. Plus, our Blog collects up-to-date analysis on major industry developments.
Outlook for 2027: toward structural AI budgets
Projections for the 2027-2028 biennium indicate that the marketing budget share allocated to generative AI tools will grow significantly. However, the growth will not be uniform. Organizations that have built solid governance and reliable metrics in 2026 will be able to scale efficiently. Conversely, those that have accumulated unvalidated implementations will have to rationalize their portfolio before they can move forward.
Therefore, 2026 is the year to build the methodological foundations. It's not yet time for widespread scaling: it's time for discipline. SMEs that invest today in defining the right metrics and governing agentive workflows are positioning themselves to reap the benefits in the following quarters.
For those who want to dive deeper into the operational implications for their organization, the team at SHM Studio can be reached here . We offer a free initial assessment to identify high-value workflows and estimate the potential ROI of agentive AI in your company's specific context.
Related articles
Discover more articles exploring similar topics, selected to offer you a more complete and stimulating perspective. Each piece of content is carefully chosen to enrich your experience.