Most AI tools were bought on a promise. Very few will be renewed on proof. That isn’t because the tools don’t work. Often they do. It’s because nobody wrote down what “working” meant before the purchase order went through. Then the renewal lands, the usage dashboard says people are logging in, and “people are logging in” quietly becomes the business case.
The renewal is usually signed off by the executive who championed the tool in the first place. They have the most context. They also have the biggest stake in the answer. That isn’t a character flaw. It’s a structural one, and it is the part of AI governance almost nobody talks about.
01Boards can already feel the gap
In PwC’s 2026 survey of about 600 US public-company directors, only 18% rated the metrics that link AI outcomes, risks and business performance as good or excellent; 66% said fair or poor. On the information they get about AI investments and return, 25% said good or excellent and 59% fair or poor.1
These are the people who approve the budget. Most of them are telling management, politely, that the evidence isn’t good enough.
PwC’s guidance to boards is blunt about why. Activity metrics such as sessions, users or tokens “may reflect experimentation, or even misuse, rather than value creation.” And “many AI costs are consumption-based rather than fixed,” so the list price in the business case is rarely the cost to run.4
02Checking changes the answer
When organizations do look properly, what they find tends to change what they run. EY surveyed 202 senior AI decision-makers at US public companies with more than $1 billion in revenue. Among those that had run a formal AI assurance review, 64% significantly modified a quarter or more of their AI systems, and 25% fully stopped a quarter or more.2
A policy isn’t proof. A review is.
Senior AI decision-makers at US public companies with $1B+ revenue, % (n=202, margin of error ±7 points)
The most common problems those reviews found were data quality, model drift and shadow AI. None of them shows up on a usage dashboard. And a policy is not the same thing as a check: 98% of the companies EY surveyed have formal AI governance policies, and 47% say they have skipped that process for an urgent deployment.2 The paper exists. The evidence often doesn’t.
Evidence also fades. In Atlassian’s survey of more than 1,100 engineers and engineering leaders, only 15% of engineers and 25% of leaders were very confident they could reconstruct the reasoning behind an AI-assisted decision six months later.3 If you can’t reconstruct a decision, you can’t show it was a good one.
The sponsor isn’t lying. They’re just the least independent person in the room.Glappy's view
03Five ways “bought” gets mistaken for “proven”
- Activity reported as results.Seats, logins and tokens show that a tool is being opened. They don’t show that anything got faster, cheaper or safer.FixOne outcome metric per tool: cycle time, cost per ticket, error rate, hours returned to a named team.
- The buyer signs off their own purchase.The person who championed the tool is the wrong person to certify the result alone.FixSign-off from someone who can say “not proven yet” without it costing them.
- The pilot ran on the easy cases.Hand-picked users and clean data make any tool look good.FixProve it on real work, with the team that will actually own it.
- There’s no “before”.Without a baseline measured before go-live, every improvement is an anecdote.FixIf you missed it, measure the manual way now, on a sample, and compare.
- Nobody totalled the cost to run.Licenses are the visible part. Consumption, integration, review time and the people who babysit the output usually aren’t.FixExpress the full cost to run per unit of outcome.
04A 45-minute test anyone can run
We wrote a scorecard for this. It works on anything with AI in it: a copilot you license, an AI feature switched on inside software you already own, an agent, an in-house model, or a pilot you haven’t bought yet. One tool, one use case, one team. Ten lines, zero to two points each. The rule that makes it work is simple: no evidence, no points.
The Proof-of-Value Scorecard
Score each line 0, 1 or 2. What a 2 (“proven”) looks like is shown under each line.
Evidence means something you could hand to a person who wasn’t in the room: a document, an export, a log, a contract clause, a dated screenshot. “I’m pretty sure” is not evidence.
Who is in the room matters as much as the questions. Invite the business owner, whoever owns IT and security for the tool, someone from finance, and one person who had nothing to do with buying it. That person holds the pen. Start every line at zero and move up only when the evidence is on screen. If the room disagrees, the lower score stands until someone produces a document.
Two lines override the total. A zero on data and access boundaries or on the human review point means you fix that first. Until it’s fixed, keep the tool away from sensitive data and have a person approve anything it produces that leaves the team.
05What to do with the number
Stop or redesign
You’re paying for hope. That’s normal for tools bought before anyone asked for evidence, and a reason not to renew on the same terms. Ask for a short term while you decide.
Prove it in 30 days
Something may be working, but you can’t show it yet. Write the proof statement first, measure a baseline, and re-score with the same people on day 30.
Scale with monitoring
You can show it works. Keep the incident log running and re-score after any model, version or pricing change.
One exception: a high score with a zero on independent sign-off is a self-assessment. Treat it as 9–14 until someone independent has checked it. The thresholds are our judgment. Adjust them to your appetite for risk, but decide them before you score, not after.
06Anyone but the buyer
Our view, which you are welcome to argue with: AI results that go into a board pack should be signed by someone who didn’t buy the tool and can say “not proven yet” without it costing them. Finance, internal audit, an outside reviewer. Anyone but the buyer.
The counterargument is fair. Sponsors understand the work in a way finance rarely does, and an outsider can miss what matters. So the sponsor stays in the room. They just don’t hold the pen. Nobody lets a vendor write its own reference. PwC makes a similar point about high-risk models, where “it may be appropriate to have independent validation or assurance of the model by internal audit or a third party.”4
“Not proven yet” is a valid result. It is much cheaper than a year of renewal on a guess.
When you want a second pair of eyes. Most teams can run the scorecard themselves, and should. It was written so you don’t need a consultant to use it. When the score is uncomfortable, when the room can’t agree, or when the real question is where AI is worth the money at all, that is the work of our AI Discovery engagement: one priority area, a half-day workshop with the people who have to live with the decision, and a written Case File with a clear next step. Run a proper pilot, set governance first, fix a data, security or infrastructure gap, or stop. Stopping is a legitimate answer. Book a time.
We hold ourselves to the standard we’re asking of you. These surveys cover large US companies, mostly public, and both PwC and EY sell AI assurance services. The Atlassian figure comes from a vendor’s survey of software teams. If you run IT for a 300-person manufacturer or a school district, your numbers will differ. The pattern is the useful part: buying is easy to show, proving is not, and hardly anyone measures the gap until a renewal forces the question.
“Bought ≠ proven”, the scoring thresholds and the independent sign-off stance are Glappy’s framework and opinion, not research findings.