Business

The underwhelming results of AI performance metrics should surprise exactly no one

Somebody at Meta built an internal leaderboard called Claudeonomics. It ranked the company’s top 250 consumers of AI tokens and handed out titles like “Token Legend” and “Cache Wizard.” Engineers set agents running for hours to climb it. In one thirty-day stretch, employees on the dashboard burned through more than 60 trillion tokens. Meanwhile, a Disney employee reportedly interacted with an AI assistant 460,000 times in nine days. JPMorgan, KPMG, Amazon, and Accenture have all been reported to be tracking employee AI activity, in some cases folding “AI-driven impact” into formal performance reviews.

The internet named the resulting behavior tokenmaxxing. People quickly learned how to make their numbers look good. They generated token usage with work that really didn’t require AI. They used agents and digital delegates to execute processes that might, or might not, be useful. The denouement came swiftly. Meta’s people took the leaderboard down once it went public. Amazon reportedly shut one down too, with some observers saying that it incentivized the employees to cheat.

Sigh. It seems that every generation or so we have to re-learn the painful lessons of previous attempts to measure progress against some brave new management idea. 

In 1956, a sociologist named V.F. Ridgway published a short paper in the very first volume of the academic journal Administrative Science Quarterly. It was literally titled “Dysfunctional Consequences of Performance Measurements.” Charles Goodhart built upon that work and formalized it in his famous law. That states that “when a measure becomes a target, it ceases to be a good measure.” We forget the painful lessons of the past and resort to the same old dysfunctions with every new management idea that captures our imaginations.

To review, 5 lessons from the past.

1. Any measure that carries consequences will be gamed

Goodhart’s Law applies not because people are dishonest, but because they are responsive. Attaching stakes to a number is an instruction, and people follow instructions. Call centers taught us this when they measured average handle time and got agents who hung up on customers. Sales organizations taught us this with quarter-end channel stuffing. Hospitals taught us this with waiting-list targets that had patients parked in ambulances outside the door.

Token counts are an unusually easy target. Gaming a sales number at least requires moving a real deal around. Gaming a token count requires very little more than a modestly inventive prompt. When the cost of manufacturing the metric approaches zero, the metric will be manufactured. The unintended consequence is that the cost of running such models is completely disassociated from the value they create.  The bill goes to someone who had nothing to do with the decision to create the cost.  One unidentified company reportedly burned through $500 million in a single month when its token usage unexpectedly ran amok. 

2. Measures that get gamed lose their variance, reducing their predictive power

A metric is only useful if it discriminates, meaning that low levels of a metric predict low levels of the thing you are measuring and high levels the opposite. If AI usage genuinely separated high performers from low ones on the day you started tracking it, that would be as intended. But the moment you attach rewards to utilization, everyone will engage in the same behavior to maximize their reward. Soon, everyone is a heavy user. The distribution flattens.

At that point the measure has no predictive power left. You cannot forecast anything from a variable that doesn’t vary. Worse, all the obvious numbers look great. The AI dashboard looks better than ever. Adoption is at 98 percent. Usage is up and to the right. Unfortunately, none of those things are providing information about whether real value is being created.

We watched this happen to performance appraisal ratings, in which the ratings themselves were “compressed” and therefore pretty useless. We watched it happen to the balanced scorecard, in which a slew of metrics designed to provide a multi-faceted view of an organization’s performance go up but leave underlying performance unchanged.

3. Inputs are not results

Tokens are an input. We keep forgetting that this is not the same as measuring outcomes.

Business Process Reengineering offers a cautionary tale. It started off with a thrilling call to action by Michael Hammer in a famous Harvard Business Review article. “Don’t automate: Obliterate” was the mandate. What was meant was to redesign obsolete practices from scratch, rather than timidly tacking on new technologies to existing processes. It’s still sound advice 36 years after the original article was published. Sprinkling a little AI fairy dust on a process from the age of COBOL isn’t going to deliver dramatic performance improvement. 

Unfortunately, in the vast majority of cases what was really incremental cost cutting got wrapped in the reengineering banner. Firms hit their targets and emerged smaller, more brittle, and no better at serving anybody. Hammer himself later acknowledged that he had underweighted the human dimension. 

4. What’s measurable crowds out what matters

The “McNamara Fallacy,” was named for Robert McNamara the U.S. Secretary of Defense during the Vietnam War. He tried to manage the war using statistical indicators like enemy body counts, but ignored data on topics such as Vietnamese resident sentiment that were difficult to measure. As Daniel Yankelovich, a sociologist, observed, it proceeds in 4 steps: 

  1. Measure what can be measured
  2. Disregard what we can’t measure
  3. Assume what cannot be measured is not important
  4. Conclude that what can’t be measured doesn’t exist

AI’s most valuable impacts are likely to be extremely difficult to measure. The value of a bad choice avoided, the benefit of quickly sorting alternatives or the increase in skill due to practicing via simulation are all hard to quantify. 

Token usage, on the other hand, is easy to measure. Measuring that can crowd out something else that isn’t showing up on any dashboard.

5. Statistical significance is not strategic relevance

Six Sigma, done well, can be extremely beneficial. Done badly, it produces rigorous, unimpeachable, verifiably significant improvements to things that do not matter. Manufacturers optimized processes for products that were already losing to substitutes. 3M’s experience under a Six Sigma-heavy regime is the case everyone cites: the operations got tighter and the innovation pipeline got thinner, and it took a successor CEO to re-ignte the messy, high-variance innovation practice the company is famous for.

If you spend enough money on AI efforts, it isn’t hard to find statistically significant operational changes. The real question is whether those numbers reflect a solid, centered, strategy for effectiveness. In too many cases, they are divorced from one another.

The questions to ask before you install another dashboard

None of this is an argument against measuring the utilization of AI. It is an argument for measuring it in the service of your strategy, not in splendid isolation. Some questions to consider:

What decision changes based on a change in the metric? If the metric doesn’t lead you to make different decisions depending on its value, what earthly purpose does it serve?

Is it an input or an outcome? Track inputs for cost control, by all means. But don’t confuse that with genuine value creation.

Does it still discriminate? When the variance in a measure collapses, the metric should be replaced.

What’s the counterweight? Every efficiency measure needs a paired quality or judgment measure that gets worse when the efficiency measure is gamed.

What assumption is this testing? The most useful early measures in genuinely uncertain territory aren’t scores at all. They are checkpoints that test specific assumptions and either validate or invalidate them.

When it comes to tokenmaxxing, management needs to understand that employees are behaving exactly as the system is designed for them to behave. There’s no reason to be surprised when rational people do things to improve the numbers on which they are being evaluated. 

We have known about how metrics can become dysfunctional for at least seven decades. It’s a shame we have to pay the same tuition for the same lessons over and over again. 

Leave a Reply

Your email address will not be published. Required fields are marked *

Are you human? Please solve:Captcha


Secret Link

Warning: foreach() argument must be of type array|object, null given in /home/misryoum/public_html/wp-content/plugins/wp-defender/src/component/class-network-cron-manager.php on line 216