Most Copilot ROI numbers
cannot be proven wrong.
That is the problem.
If you are being asked to justify the Copilot renewal, you have probably met the standard toolkit: a calculator, three scenarios labelled conservative, moderate and aggressive, and a headline percentage. A CFO who has seen a few of these knows exactly how much they are worth.
This page is the honest version. What the widely quoted numbers actually are, why self-reported time savings are the weakest evidence you can bring, and the small number of things you can measure that hold up in a real budget conversation.
Where the famous number comes from
The 116% figure is from a Forrester Total Economic Impact study commissioned by Microsoft. It models a composite global enterprise with 25,000 employees and 6.25 billion dollars in revenue, and arrives at 36.8 million dollars of benefit against 17.1 million dollars of cost over 3 years, with a 10-month payback.
Two things are true at once. TEI studies are a legitimate, well-understood modelling format, and Forrester is transparent about who commissioned them. And a model of a company that does not exist, paid for by the vendor, is not evidence about your company.
The 9 hours a month problem
The same study reports employees saving an average of 9 hours per month. That figure travels everywhere, and it is worth being precise about what it is: people estimating how much time they think they saved.
Self-reported savings are the weakest evidence class there is. People are generous about tools they like, generous about anything they were just trained on, and genuinely bad at estimating time. Ask the same group how long their weekly report used to take and you will get a wide range for a task they have done 40 times.
That does not make it worthless. It makes it a sentiment measure, and it should be labelled as one when it reaches a slide.
The measurement you skipped
Here is the uncomfortable part. Almost nobody measures anything before the rollout, because before the rollout everyone is busy rolling out. Then 6 months later someone asks for proof, and there is no before to compare the after against.
Without a baseline, every number you produce is an anecdote with a percentage sign on it. This is not a Copilot flaw. It is the single most common reason a genuinely successful rollout cannot defend itself at renewal.
What actually holds up in a budget meeting
Fewer metrics, better chosen. These are the ones I have seen survive contact with a finance team.
01 Active use, not licenses assigned
The share of licensed people who used it in the last 30 days. This one is not a value metric, it is a sanity metric, and it comes first because every other number is meaningless if this one is low. Benchmark: about 36% of people with access actually use it.
02 One task, timed before and after
Pick something specific and recurring: the monthly report, the customer response, the meeting summary. Time it for a few people before, time it after, state the sample size honestly. One credible measured task beats a spreadsheet of estimates.
03 Throughput on something the business already counts
If a team already tracks tickets closed, proposals sent, or documents reviewed, that number exists whether or not anyone is doing an AI project. It is much harder to argue with, because nobody created it to win this argument.
04 What stopped happening
Work that used to get outsourced, a queue that no longer backs up, a report that no longer needs a person on a Sunday. This is often the realest saving and the one nobody writes down.
What I tell clients who want a big number
Sometimes the honest answer is that it is too early. If usage is at 20%, there is no ROI question yet, there is an adoption question, and measuring harder will not fix it. That conversation is on the why rollouts stall page.
And occasionally the honest answer is that this team does not need the licenses. I would rather say that than help build a number that falls apart the first time a CFO reads the methodology. It is also, in my experience, the thing that makes people trust the rest of the numbers.
Questions that come up
Is the Forrester study useless then?
No. It is a reasonable model and it is useful for framing the shape of the investment. It is just not measurement of your organization, and the distinction matters when someone senior asks where the number came from.
What is a realistic payback period?
It depends almost entirely on how many people actually use it, which is why the usage number comes first. I would be cautious about anyone quoting you a specific month without having looked at your usage data.
We never took a baseline. Is it too late?
Not entirely. You can baseline a task that has not been touched by Copilot yet, or run a controlled comparison between a trained team and an untrained one. Less clean than doing it up front, and far better than nothing.
Does the Copilot usage dashboard prove value?
It proves usage, which is a necessary condition and not a sufficient one. Interactions per week is an activity count, not an outcome.
Can you help build the business case?
I can help you decide what is worth measuring and interpret what comes back. I am not going to hand you a calculator that produces a number you cannot defend.
Need a number you can actually defend?
An intro call takes 30 to 45 minutes. We look at what you can realistically measure from where you are now, and whether the honest answer is a business case or a different conversation about adoption.