Most Copilot ROI numbers
cannot be proven wrong.
That is the problem.
If you are being asked to justify the Copilot renewal, you have probably met the standard toolkit: a calculator, 3 scenarios labelled conservative, moderate and aggressive, and a headline percentage. A CFO who has seen a few of these knows exactly how much they are worth.
This page is the honest version. What the widely quoted numbers actually are, why self-reported time savings are the weakest evidence you can bring, and the small number of things you can measure that hold up in a real budget conversation.
Where the famous number comes from
The 116% figure is from a Forrester Total Economic Impact study commissioned by Microsoft. It models a composite global enterprise with 25,000 employees and 6.25 billion dollars in revenue, and arrives at 36.8 million dollars of benefit against 17.1 million dollars of cost over 3 years, with a 10-month payback.
Two things are true at once. TEI studies are a legitimate, well-understood modelling format, and Forrester is transparent about who commissioned them. And a model of a company that does not exist, paid for by the vendor, is not evidence about your company.
The 9 hours a month problem
The same study reports employees saving an average of 9 hours per month. That figure travels everywhere, so be precise about what it is: people estimating how much time they think they saved.
Self-reported savings are the weakest evidence class there is. People are generous about tools they like, generous about anything they were just trained on, and genuinely bad at estimating time. Ask the same group how long their weekly report used to take and you will get a wide range for a task they have done 40 times.
That does not make it worthless. It makes it a sentiment measure, and it should be labelled as one when it reaches a slide.
The measurement you skipped
Here is the uncomfortable part. Almost nobody measures anything before the rollout, because before the rollout everyone is busy rolling out. Then 6 months later someone asks for proof, and there is no before to compare the after against.
Without a baseline, every number you produce is an anecdote with a percentage sign on it. And it is the most common reason a genuinely successful rollout cannot defend itself at renewal (Copilot has nothing to do with it).
What actually holds up in a budget meeting
Fewer metrics, better chosen. These are the ones I have seen survive contact with a finance team.
01 Active use, not licenses assigned
The share of licensed people who used it in the last 30 days. Call it a sanity metric. It comes first because every other number is meaningless if this one is low. There is no reliable public benchmark for this one. The closest is Recon Analytics (February 2026): only 35.8% of US workers who have Copilot at work use it as their main AI tool. That measures something narrower than 30-day use, so your own baseline is the real benchmark.
02 One task, timed before and after
Pick something specific and recurring: the monthly report, the customer response, the meeting summary. Time it for a few people before, time it after, state the sample size honestly. One credible measured task beats a spreadsheet of estimates.
03 Throughput on something the business already counts
If a team already tracks tickets closed, proposals sent, or documents reviewed, that number exists whether or not anyone is doing an AI project. It is much harder to argue with, because nobody created it to win this argument.
04 What stopped happening
Work that used to get outsourced, a queue that no longer backs up, a report that no longer needs a person on a Sunday. This is often the realest saving and the one nobody writes down.
What I tell clients who want a big number
Sometimes the honest answer is that it is too early. If usage is at 20%, you do not have an ROI question yet. You have an adoption question, and measuring harder will not fix it. That conversation is on the why rollouts stall page.
And occasionally the honest answer is that this team does not need the licenses. I would rather say that than help build a number that falls apart the first time a CFO reads the methodology. It is also, in my experience, the thing that makes people trust the rest of the numbers.
Questions that come up
Is the Forrester study useless then?
No. It is a reasonable model and it is useful for framing the shape of the investment. It is just not measurement of your organization, and the distinction matters when someone senior asks where the number came from.
What is a realistic payback period?
It depends almost entirely on how many people actually use it, which is why the usage number comes first. I would be cautious about anyone quoting you a specific month without having looked at your usage data.
We never took a baseline. Is it too late?
Not entirely. You can baseline a task that has not been touched by Copilot yet, or run a controlled comparison between a trained team and an untrained one. Less clean than doing it up front, and far better than nothing.
Does the Copilot usage dashboard prove value?
It proves usage, which is a necessary condition and not a sufficient one. Interactions per week is an activity count, not an outcome.
Can you help build the business case?
I can help you decide what is worth measuring and interpret what comes back. I am not going to hand you a calculator that produces a number you cannot defend.
Need a number you can actually defend?
An intro call takes 30 to 45 minutes. We look at what you can realistically measure from where you are now, and whether the honest answer is a business case or a different conversation about adoption.