Wednesday, September 30, 2026

Today’s Edition

AI Intel Report

MARKETS —

Enterprise AI

OpenAI, David Sacks Push AI Scorecards to Prove ROI and Net Job Gains

A new measurement push from OpenAI and David Sacks aims to tie AI spending to completed work, full unit costs and net hiring, giving CFOs an evidence-based alternative to model benchmarks.

14 MIN READ
Finance leader points at an upward trend chart during an AI ROI meeting in a city office.
Illustration: AI Intel Report

An enterprise AI scorecard is a structured measurement framework that ties AI spending to completed work, full cost per successful task, trustworthiness and value at scale.

OpenAI and David Sacks, the former White House AI and crypto czar, are pressing enterprises to replace anecdotal justifications for AI spending with empirical scorecards that measure investor returns, productivity and worker outcomes. The push arrives as finance chiefs and chief AI officers face mounting pressure to show that model procurement converts into completed work, lower unit costs and net hiring rather than displacement. Two parallel efforts define the moment: OpenAI's proposed 'Useful Intelligence per Dollar' framework and Sacks' promotion of firm-level and occupation-level studies tying AI adoption to employment and wage gains.

Both efforts treat measurement as the missing discipline in enterprise AI. Model benchmarks such as MMLU or leaderboard scores describe capability, not business value, and pilot programs often stop at qualitative feedback. The new scorecard proposals attempt to link dollars spent to tasks successfully completed, including the hidden costs of retries, human review, security, privacy and governance. If adopted widely, they could change how AI budgets are approved, how vendors are evaluated and how labor-market effects are reported.

Why do enterprises need an AI scorecard now?

Enterprise AI budgets have expanded rapidly, and so has scrutiny. Procurement teams now ask vendors to define what a 'successful task' means before a deployment, rather than after. Chief financial officers want line items that separate model usage from employee time, rework and oversight. The scorecard movement responds to a gap in standard accounting: traditional ROI calculations capture token costs but routinely omit the labor and governance overhead required to make AI output reliable.

Labor-market anxiety adds urgency. Public debate has centered on whether generative AI displaces workers, and vendors have struggled to rebut those fears with data. Sacks has used recent studies to argue that AI-exposed occupations are adding jobs and raising real wages, calling displacement fears a 'hoax.' The claim is contested, but it has shifted attention to measurement: if adoption produces net hiring, that outcome should be visible in payroll records, not asserted in marketing materials.

What is OpenAI's 'Useful Intelligence per Dollar' scorecard?

OpenAI introduced the measurement framework in a post titled 'AI 時代のスコアカード' — 'A Scorecard for the AI Age' — on its Japanese-language site. The proposed metric is 'useful intelligence per dollar,' which OpenAI frames as the ultimate scorecard for AI spending. The framework answers four questions: Does AI complete important work? What does one successful task cost? Can people trust the results? Does each additional dollar create more value as usage grows?

Pulse2, which summarized the framework, reported that the measurement considers the complete cost of achieving a usable outcome. That includes model usage, computing resources, employee time, human review, retries and any rework. Security, privacy, safety and governance remain part of the economic calculation rather than being treated as separate compliance costs. The design is meant to correct the common error of comparing sticker-price token rates without accounting for the labor needed to make AI output usable.

  1. Useful work completed: the volume of accepted deliverables produced by AI in a defined period.
  2. Full cost per successful task: tokens, compute, employee time, human review, retries, rework and governance overhead divided by accepted tasks.
  3. Dependability and trust: human acceptance rates, error rates, escalation frequency and audit outcomes.
  4. Value at scale: marginal cost per task and revenue or productivity per AI dollar as usage grows.
OpenAI's proposed scorecard dimensions and the questions they answer
Scorecard dimensionCore questionWhat enterprises should track
Useful work completedDoes AI finish important jobs?Number of accepted deliverables per week, completion rate
Full cost per successful taskWhat does success cost, all-in?Tokens, compute, employee time, human review, retries, security and governance overhead
Dependability and trustCan people rely on the output?Human acceptance rate, error rate, escalation frequency
Value at scaleDoes each extra dollar create more value?Marginal cost per task, usage growth, revenue per AI dollar

How does a 'useful intelligence per dollar' scorecard differ from existing AI ROI metrics?

Existing AI ROI metrics tend to fall into two categories. The first is model-centric: latency, throughput, benchmark scores and cost per million tokens. The second is outcome-centric but vague: productivity gains, time saved or qualitative feedback from pilot users. OpenAI's scorecard sits between them. It demands a measurable unit of useful work, a full cost per successful task and a trust metric, which means the model's technical performance is only one input into a business calculation.

Pulse2 noted that the framework considers the complete cost of achieving a usable outcome, including employee time and human review. That is a meaningful shift. A cost per million tokens is a procurement metric; a cost per successful task is an operating metric. The former tells a buyer what the model costs to run, while the latter tells a CFO what the business pays to get a job done. Enterprise AI leaders are increasingly being asked for the second number.

How does the Ramp and Revelio study measure AI's impact on jobs?

The most detailed evidence cited by Sacks comes from Ramp Economics Lab and Revelio Labs. They linked observed AI spending from Ramp card and bill pay data to Revelio Labs workforce records for 21,559 firms in the United States. The study compares firms in the top third of AI spending per employee, defined as high-intensity adopters, with firms making smaller investments. It is one of the largest firm-level analyses of generative AI spending and employment to date.

The results run counter to displacement predictions. High-intensity AI adopters grew employment by 10.2% over two years after adoption, while low-intensity adopters showed no statistically significant change. Entry-level headcount rose 12% for high-intensity adopters. The findings held across roles including engineering, sales, administration and customer service, with gains concentrated in the Information sector.

The gains were not immediate. Employment effects emerged gradually six to 12 months after adoption, a timeline that matters for evaluation design. A 30-day scorecard will capture early productivity signals but will likely miss the full labor-market response. Enterprises that halt a deployment after one month may therefore understate the long-term return, while those that measure only token costs may overstate it.

What does David Sacks say the labor data shows?

Sacks, who served as White House AI and crypto czar and now co-chairs the President's Council of Advisors on Science and Technology, has amplified the Ramp and Revelio findings. In an X post citing the study, he wrote that companies that adopt AI tend to grow faster following adoption, and that the results counter predictions that AI adoption will lead to broad job loss. His post framed the firm-level data as direct evidence against displacement narratives.

He has also cited occupation-level research showing that occupations with high AI exposure grew jobs by 1.7% versus 0.8% for other occupations, and real wages by 3.8% versus 0.7%. Sacks described displacement fears as a 'hoax' in that context. The statistics come from a Vanguard study covered by Benzinga, and they are now circulating as a rejoinder to forecasts of AI-driven unemployment.

AI has dramatically lowered the cost of writing code... there is now far more code to manage than ever before... a massive productivity boom driven by the proliferation of bespoke software throughout the entire economy.David Sacks, former White House AI and Crypto Czar

Sacks' broader argument connects labor outcomes to software economics. AI, he argues, has dramatically lowered the cost of writing code, which expands the volume of software that organizations can justify building. More software means more code to manage, more bespoke applications across the economy and more demand for the engineers, product managers and operators who support them. The result, in his view, is a productivity boom rather than mass unemployment.

What does a 30-day bounded workflow evaluation look like?

OpenAI's scorecard framework includes a bounded workflow evaluation designed to produce evidence within 30 days. The method requires selecting a specific workflow, pre-agreeing acceptance criteria for what counts as a successful task, and then measuring outcomes against those thresholds before any enterprise-wide rollout. This bounded design addresses a common failure mode: organizations measure model performance in isolation and never connect it to the cost and quality of the business process it serves.

  1. Select one bounded workflow with measurable outputs, such as customer support ticket resolution or invoice processing.
  2. Define acceptance criteria before deployment, including quality thresholds and error tolerances.
  3. Capture full costs, including employee time, human review, retries and security or governance overhead.
  4. Run the evaluation for 30 days, tracking dependability metrics alongside cost.
  5. Compare results to the pre-agreed thresholds and decide whether to scale, adjust or halt.

Acceptance criteria are the hardest part of the process. A 'successful task' cannot be defined as any output the model produces; it must meet a standard that a human expert would accept. For customer support, that might mean a resolved ticket with no escalation. For invoice processing, it might mean data extraction that matches a human audit with zero critical errors. Pre-agreeing these standards prevents teams from redefining success after the fact.

How do security, privacy and governance costs enter the calculation?

OpenAI's framework explicitly includes security, privacy, safety and governance as part of the economic calculation. That is a notable departure from earlier ROI models, which treated compliance as an external constraint. In practice, these costs include prompt-injection testing, data-loss prevention, model access controls, audit logging, human review for high-risk outputs and legal review of vendor terms. Each category adds labor and infrastructure that must be divided by the number of successful tasks.

The inclusion of these costs changes procurement math. A model with a higher token price but lower governance overhead can be cheaper per successful task than a lower-priced model that requires extensive human review. Conversely, a model that appears inexpensive on a per-token basis may fail acceptance criteria frequently, driving up retry and review costs. The scorecard is designed to surface those tradeoffs instead of hiding them in an IT budget.

What data do enterprises need to build their own scorecard?

Building a scorecard requires data from at least four systems. Procurement and finance systems supply AI spend per vendor and per business unit. Application logs supply usage volumes, retries, error rates and human escalations. HR systems supply the employee time spent on review, rework and supervision. Governance teams supply the cost of security reviews, privacy assessments and audit activities. Most enterprises have these data in separate silos, and the scorecard is an argument for integrating them.

Ramp and Revelio's methodology is a model for that integration. They combined observed spending data with workforce records to measure employment changes, which is exactly the kind of cross-functional data merge that an enterprise scorecard requires. A company that cannot link its AI spend to its HR and operations data will struggle to produce credible ROI evidence, regardless of how sophisticated its models are.

What does the evidence mean for entry-level hiring?

The entry-level finding is one of the most consequential in the Ramp and Revelio study. High-intensity adopters saw entry-level headcount rise 12%, a figure that directly challenges the fear that AI will eliminate the first rung of the career ladder. One interpretation is that AI automates routine tasks and makes junior employees more productive, allowing firms to hire more of them. Another is that high-growth firms both adopt AI and expand hiring, with no direct causal link.

Either way, the data gives talent leaders a new baseline. If entry-level hiring rises after AI adoption, then workforce planning should assume expansion, not contraction. But the averages also conceal variation by sector and role; gains were concentrated in the Information sector, and administrative roles may see different effects. A scorecard that tracks employment by level and function would give each enterprise its own evidence rather than relying on aggregate studies.

What are the market and investor implications?

If scorecards become standard, the competitive advantage shifts from model capability claims to deployment discipline. Enterprises that adopt pre-agreed acceptance criteria and full-cost accounting will be able to show boards and investors a defensible return on AI spending. Vendors that publish comparable scorecards may earn procurement preference, while vendors that resist measurement may face slower enterprise adoption.

Investors are a primary audience for the new data. Sacks has cited Vanguard research showing AI-exposed occupations gained jobs and wages, and the Ramp and Revelio data suggests adoption correlates with firm growth. For public markets, these findings support the thesis that AI spending produces top-line expansion rather than margin compression. The scorecard movement gives investors a framework for distinguishing AI investments that generate measurable output from those that merely accumulate compute costs.

How should vendors respond to the scorecard movement?

Vendors face a strategic choice. They can publish their own scorecards, using customer deployments to show useful work completed, full cost per successful task and dependability metrics. Or they can continue selling on model capability claims and token prices, which exposes them to procurement teams that now demand operating metrics. The early movers are likely to be vendors that can instrument their platforms to report these numbers automatically.

OpenAI's own framework is a signal that model providers intend to lead on measurement rather than resist it. By defining the scorecard, OpenAI shapes how enterprises calculate value and which costs they include. That has commercial implications: a framework that includes security, privacy and governance overhead favors vendors that can demonstrate lower governance costs, not just lower token prices.

What are the limitations of the new scorecard approaches?

The new frameworks have real limitations. The Ramp and Revelio study measures correlation between observed AI spending and employment changes, not a controlled experiment; high-intensity adopters may differ in other ways, such as access to capital or digital maturity. The occupation-level data cited by Sacks captures exposure to AI, not actual deployment per worker, and averages can hide concentrated losses in specific roles or regions.

Scorecard design also invites gaming. If acceptance criteria are set by the team that operates the AI system, teams may lower thresholds to report success. If security and governance overhead is excluded, the full cost per successful task is understated. OpenAI's framework explicitly includes these overhead categories, but the incentives to undercount them remain. Independent audit or a second-line review function is needed to keep the scorecard credible.

The 30-day evaluation window may also miss long-term effects. The Ramp and Revelio data found employment gains emerged six to 12 months after adoption, suggesting that short evaluations can understate both costs and benefits. A disciplined organization should treat the 30-day scorecard as a gate, not a final verdict, and re-run the measurement at quarterly intervals as usage grows and workflows mature.

What are the policy implications of the scorecard push?

The scorecard movement also has a policy dimension. Sacks co-chairs the President's Council of Advisors on Science and Technology, and his public promotion of firm-level and occupation-level data is part of a broader argument that AI regulation should not be premised on mass displacement. If scorecards become a standard way of reporting AI's labor effects, policymakers gain an evidence base for decisions about training, unemployment insurance and AI adoption incentives.

The proposed framework could also influence disclosure. If large enterprises adopt scorecards, investors and regulators may begin to expect them, much as they expect carbon accounting or diversity reporting. That would be a significant change from today's environment, in which most companies disclose AI spending only in aggregate and rarely report cost per successful task or employment effects.

What should enterprises do next?

Enterprises should treat the scorecard as a governance artifact, not a marketing metric. The practical first step is to inventory existing AI deployments and identify which have defined success criteria and full-cost data. Most organizations will find that they measure tokens and latency but not cost per successful task or human acceptance rates. That gap is the starting point for the new framework.

Finance and AI leaders should jointly define acceptance criteria for one high-value workflow, run a 30-day bounded evaluation, and publish the results internally. The evaluation should include security, privacy and governance overhead from the start, because excluding those costs produces an artificially favorable return. If the results clear the pre-agreed thresholds, the organization can expand with evidence; if not, it can stop before sunk costs grow.

Labor outcomes should also be tracked with the same rigor. The Ramp and Revelio methodology shows that payroll data can be linked to AI spending to measure net hiring effects. Enterprises that want to counter displacement fears should measure entry-level headcount and role-level employment changes in their own HR and procurement records, and report them alongside financial returns. The scorecard, in other words, is not only a CFO tool; it is a workforce policy instrument.

How does this fit with the broader enterprise AI cycle?

The scorecard push arrives at a specific point in the enterprise AI cycle. The first phase was experimentation, funded by fear of missing out. The second phase is consolidation, in which CFOs demand evidence that pilots produce returns. The third phase, which these proposals represent, is standardization of measurement. OpenAI and Sacks are trying to define the vocabulary of that phase: useful intelligence per dollar, successful tasks, full costs, dependability and net employment effects.

Enterprises that adopt the vocabulary early will have an advantage in budgeting, vendor negotiations and workforce communications. Those that wait may find themselves defending AI programs without the data to prove value. The core message of the scorecard movement is simple: measure the work AI completes, the full cost of that work, the trustworthiness of the output and the effect on headcount, and the debate over AI's value will become an accounting exercise rather than a marketing war.

Frequently asked

What is the 'Useful Intelligence per Dollar' scorecard?

OpenAI proposed it as a framework for measuring enterprise AI value. It answers four questions: whether AI completes important work, what a successful task costs, whether people trust the results, and whether each additional dollar creates more value as usage grows.

Do high-intensity AI adopters add jobs or cut them?

A Ramp Economics Lab and Revelio Labs study of 21,559 U.S. firms found high-intensity adopters grew employment by 10.2% over two years after adoption, with entry-level headcount rising 12%. Low-intensity adopters showed no statistically significant change.

How can an enterprise test AI returns in 30 days?

Select one bounded workflow, define acceptance criteria before deployment, capture full costs including human review and governance overhead, and measure dependability and cost for 30 days. Compare results against the pre-agreed thresholds before scaling.

What did David Sacks say about AI job displacement?

He called displacement fears a 'hoax,' citing data showing AI-exposed occupations grew jobs 1.7% versus 0.8% for others and real wages 3.8% versus 0.7%. He also argued that lower coding costs expand demand for software and the people who manage it.

Sources

  1. Ramp Economics Lab / Revelio Labs — High-intensity AI adopters grew employment by roughly 10.2% following adoption, with entry-level headcount rising 12%; low-intensity adopters showed no statistically significant change.
  2. OpenAI — OpenAI proposed the 'Useful Intelligence per Dollar' scorecard, answering four questions about useful work, cost per successful task, trust and value at scale.
  3. Pulse2 — The framework considers the complete cost of achieving a usable outcome, including model usage, computing resources, employee time, human review, retries, rework, and security, privacy, safety and governance overhead.
  4. Benzinga — AI-exposed occupations grew jobs 1.7% versus 0.8% for others, and real wages 3.8% versus 0.7%.
  5. Benzinga — Sacks argued that AI has lowered the cost of writing code, creating a productivity boom driven by bespoke software proliferation.
  6. X (David Sacks) — Sacks cited the Ramp and Revelio study and wrote that companies that adopt AI tend to grow faster following adoption, countering predictions of broad job loss.