- The problem that the scorecard aims to solve
- The four dimensions: framework architecture
- 1. Useful Work: the work that really matters
- 2. Cost per Successful Task: the real cost of autonomy
- 3. Dependability: reliability as a strategic asset
- 4. Return on Compute: infrastructure efficiency
- Concrete applicability for Italian SMEs
- Trade-offs to consider before adopting the framework
- What the metrics don't say
- The recommended decision: where to start
Sarah Friar, CFO of OpenAI, published a operational scorecard for the AI era . The tool proposes four key metrics: useful work , cost per successful task , dependability and return on compute . Therefore, the debate on AI ROI stops being abstract and becomes measurable.
However, many Italian companies — including SMEs and mid-market — continue to evaluate AI with metrics inherited from traditional software: licenses, hours saved, costs avoided. These measures are insufficient. In fact, an AI agent that completes 60% of tasks autonomously but fails on the critical 40% produces a negative ROI, regardless of the declared hourly savings. OpenAI's scorecard introduces a paradigm shift: it measures the useful work actually completed , not the theoretically saved time.
At SHM Studio we follow this framework closely. Likewise, we apply it to projects of AI integration that we support for our clients. In summary, this article analyzes the four dimensions of the scorecard, evaluates its practical applicability for Italian companies, and indicates the trade-offs to consider before adopting it as a decision-making compass.
The problem that the scorecard aims to solve
For years, the ROI of artificial intelligence has been measured with approximations. Companies counted saved hours, estimated avoided costs, and multiplied by the average hourly rate. The result was a reassuring number, often far from operational reality.
However, with the spread of autonomous AI agents , this approach becomes dangerous. An agent working autonomously on critical processes — from order management to customer care — cannot be evaluated solely on the theoretical time saved. Therefore, a metric is needed to measure the value actually produced .
Sarah Friar, CFO of OpenAI, tackled this problem head-on. In a paper published on OpenAI's official website, she introduced a practical scorecard for the AI era . The framework is structured into four dimensions. Each answers a precise question about the value generated by artificial intelligence in real business contexts.
The four dimensions: framework architecture
First of all, it is helpful to understand the overall logic. The four metrics are not independent. They form a coherent system that measures AI along the axis of value produced, cost incurred, reliability, and computational efficiency.
1. Useful Work: the work that really matters
The first dimension is the useful work . This is not about measuring how many requests the AI has processed, but how many it has completed with a useful outcome for the business. In fact, a chatbot that answers a thousand questions but solves only two hundred problems has a useful work rate of 20%.
This metric requires companies to define in advance what "useful completion" means. For example, in a lead generation context, a useful task might be correctly qualifying a contact. In a technical support context, it could be resolving a ticket independently without human escalation.
Consequently, useful work forces marketing and operations teams to clarify their goals even before implementing any AI solution. It is an exercise in strategic clarity, not just technical. For this reason, we at SHM Studio we consider it the starting point of any project of Digital marketing with an AI component.
2. Cost per Successful Task: the real cost of autonomy
The second dimension is the cost per successful task . This metric calculates how much every successfully completed task costs the company. Not the total cost of the AI system, but the unit cost of the useful result.
However, the calculation is less intuitive than it seems. It includes the cost of APIs or the model, the cost of failure (uncompleted tasks requiring human intervention), the cost of supervision, and the cost of error correction. Therefore, a seemingly inexpensive system can turn out to be costly if the failure rate is high.
According to recent research by McKinsey on the economic potential of generative AI , companies that measure the cost per useful output — rather than total system cost — get a 40-60% more accurate picture of actual ROI. Therefore, the metric is not just theoretical: it has a direct impact on budget decisions.
3. Dependability: reliability as a strategic asset
The third dimension is the dependability . It measures the AI's ability to behave predictably and reliably over time. It's not enough for the system to work well on average. It needs to work well even in edge cases, during peak loads, and in contexts not foreseen during training.
Despite this, dependability is often the most overlooked dimension in preliminary evaluations. Demos always work. Production systems, however, encounter real variability. In fact, an AI agent managing Google Ads campaigns autonomously must be reliable even during peak seasons, not just during ordinary periods.
For Italian companies operating in sectors with strong seasonality — retail, tourism, food — this dimension is critical. We at SHM Studio we systematically evaluate it in projects that integrate AI into google ads campaigns and in the LinkedIn campaigns .
4. Return on Compute: infrastructure efficiency
The fourth dimension is the return on compute . It measures the value produced for each unit of computational power used. It is a more technical metric, but also relevant for SMEs using pay-as-you-go cloud services.
Furthermore, with AI model costs constantly evolving, return on compute becomes an indicator of medium-term economic sustainability. A system that currently delivers adequate value could become inefficient if computational costs rise or if more efficient models become available.
Furthermore, this metric incentivizes companies to choose the right model for the right task. The most powerful model is not always the most efficient. For example, for repetitive and well-defined tasks, lighter models can offer a significantly higher return on compute.
Concrete applicability for Italian SMEs
Friar's framework was born in an enterprise context. However, its logic can also be applied to Italian SMBs and the mid-market, with some adaptations.
Specifically, SMBs rarely have granular monitoring systems. Therefore, implementing the scorecard requires a preliminary investment in data infrastructure . Without structured AI task logs, outcome tracking, and alerting systems, the four metrics remain out of reach.
However, this doesn't mean the scorecard is useless for SMEs. On the contrary, it can guide the choice of AI tools from the very beginning. For example, a company evaluating a chatbot for customer support should ask the provider: how do I measure useful work? How do I calculate the cost per successful task? These are the correct selection criteria, not the number of available integrations or the graphical interface.
According to Gartner AI Trends Report , by 2027 over 60% of mid-market companies will adopt structured frameworks for measuring AI ROI. As a result, those who start building this measurement capability now gain a concrete competitive advantage.
For marketing managers, the framework translates into immediate operational questions. How many of the AI automations active today produce measurable useful work? What is the cost per qualified lead generated with AI support compared to the manual process? The SEO Strategy AI-supported content produces higher performance? These questions drive more solid budget decisions.
Trade-offs to consider before adopting the framework
The OpenAI scorecard is a powerful tool. However, it comes with some trade-offs that decision-makers must consider before adopting it.
The first trade-off concerns the implementation complexity . Measuring the four dimensions requires technical infrastructure, logging processes, and analytical skills. For less structured organizations, the cost of implementing the measurement system may, in the short term, outweigh the informational benefits obtained.
The second trade-off concerns the definition of «success» . Useful work and the cost per successful task depend on a clear definition of what constitutes a positive outcome. In complex contexts — such as SEO content production or managing multi-channel campaigns — this definition is often ambiguous and debated internally.
The third trade-off concerns the risk of local optimization . Optimizing the four metrics individually can produce suboptimal system-level behaviors. For example, increasing dependability by reducing edge cases handled by AI can lower overall useful work. Therefore, metrics should be read in an integrated, not isolated, manner.
Finally, the framework measures the value of AI quantitatively. However, some AI benefits—such as perceived interaction quality or brand consistency in communication—are hard to quantify with these metrics. Therefore, the scorecard should be combined with qualitative assessments.
What the metrics don't say
There is a dimension that Friar's framework does not explicitly capture: the cost of inertia . Companies that don't adopt AI — or adopt it without measuring it — don't have a zero cost. They have a growing opportunity cost.
According to Harvard Business Review , organizations that develop structured AI measurement capabilities achieve 2.5x higher returns than those that adopt AI without an evaluation framework. Therefore, the scorecard is not just a control tool. It is an enabler of organizational learning.
Furthermore, the framework pushes companies to build a measurement culture around AI. This culture is transferable: teams that learn to measure the ROI of an AI agent apply the same logic to subsequent projects, reducing learning cycles and accelerating time-to-value.
For marketing and digital managers of Italian companies, this means that investing in the measurement framework today — even if AI is still experimental — builds an organizational competency that will become increasingly relevant in the coming years. Services from web development and of AI integration that we follow at SHM Studio always start from this premise: measure first, scale later.
The recommended decision: where to start
For Italian companies that want to adopt OpenAI's scorecard, we suggest a three-phase approach.
- Phase 1 — Mapping: identify active AI processes and define what constitutes a «successfully completed task» for each. This exercise requires involvement from both the technical and business teams.
- Phase 2 — Instrumentation: implement structured logging of AI outcomes. Without data, metrics remain theoretical. Even simple solutions — shared spreadsheets, Looker Studio dashboards — are a valid starting point.
- Phase 3 — Periodic review: establish a review cadence for the four metrics. Monthly for high-volume systems, quarterly for more stable systems. The scorecard is only useful if it is read and acts as a decision-making input.
Therefore, adopting the framework doesn't necessarily require a large initial investment. It requires methodological discipline and clarity on objectives. These are prerequisites that any organization can develop, regardless of size.
To learn more about how to apply these principles to projects of Digital marketing and Artificial intelligence of your company, the SHM Studio team is available for an initial consultation. You can contact us through the contact page or explore the other insights in SHM Studio blog .
Related articles
Discover more articles exploring similar topics, selected to offer you a more complete and stimulating perspective. Each piece of content is carefully chosen to enrich your experience.