Copilot ROI for Management: What to Measure, and What to Stop Measuring
Bottom line: licences bought is not a metric. Logins is not a metric. The 2 numbers worth reporting to leadership are weekly active usage 3 months after training, and what happened to one named process. Everything else is activity data dressed up as results.
Every organisation I work with wants to know the ROI. Very few of them have set up anything that could answer the question.
And I understand why. Measuring this properly means committing to a number before you know whether you will like it.
I spend most of my time with senior leadership teams, and I will admit I am a sucker for those rooms. It is maximum impact. They are the only group that can require an entire organisation to learn this, and the only group that can decide real time and budget go into it. When leadership understands it deeply the organisation moves, and when it does not, everything stays at pilot level and there is nothing worth measuring in the first place. The rooms I enjoy most are the ones asking how, not whether, because those are the ones that end up with a number at all.
The number that should worry you
BCG surveyed 152 CEOs in July 2026 at companies above $500M in revenue. 88% reported a benefit in cost or revenue somewhere. 26% had turned it into a broad business change. And 14% could say what it did to the profit and loss.
PwC surveyed 351 CEOs across 59 countries in August 2026. Only 39% said AI had already increased revenue or cut costs for them. And 20% of companies are capturing 74% of all the economic value being created.
Put those together and the picture is blunt. Almost everyone reports a benefit. Almost nobody measures it. And most of the actual value is going to a small group who figured this out.
The measurement is not paperwork. It is the thing that puts you in the 20%.
Stop reporting these
Licences purchased. This is a procurement number. It tells you what you spent, not what happened.
People who logged in at least once. I see this in board packs constantly. A report saying 200 people opened it hides the problem instead of exposing it.
Training completion. Attendance is not capability. I run the training and I will tell you plainly: people sitting through my session proves nothing on its own.
Enthusiasm in the room. Every workshop ends with energy. That energy has a half-life of about 3 weeks without follow-up. Do not put it in a report.
Time saved, self-reported. Ask 50 people how much time they saved and you will get a number. It will be wrong in both directions, and you cannot build a case on it.
Measure 2 things instead
1. Weekly active usage, at the 3 month mark
How many people used it in a given week, 3 months after the training, with nobody reminding them.
3 months matters. Anything measured in week 2 is measuring novelty. The interesting question is whether it survived contact with a normal busy month.
Useful benchmark to hold in your head: market average adoption sits around 35.8%, which means roughly 2 thirds of the licences organisations pay for never become a habit. If you are above that, something is working. If you are at 12%, you have a management problem, not a tool problem.
2. Business impact on one named process
This is where most people give up, because it feels hard. It is not hard. It is just specific.
Pick the process you named when you approved the budget. Monthly reporting pack. Contract review cycle. Tender response. Customer complaint analysis. Measure the same thing you measured before: elapsed days, number of people involved, or rework rate.
One process, measured properly, beats a company-wide productivity estimate that nobody believes.
How to set the KPI so it survives contact with a CFO
Make it quantitative and dated. Not "increase usage". Something like: 60% of managers bring an analysis prepared with Copilot to the monthly review by the end of the quarter.
Give it an owner from the executive team, by name. The CEO and the C-suite own this project, and the KPI should sit against one of them rather than against a committee.
Set the baseline before you start. This is the step everyone skips and then regrets, because 6 months later nobody can remember how long the process used to take, and the conversation collapses into opinion.
And write down what would make you stop. A KPI you would never fail is not a KPI.
The costs people forget to count
On the other side of the ratio, licences are usually the small and predictable part. A few tens of dollars per user per month.
What does not get costed: the training hours, the time of whoever runs this internally, the champion time (3 to 4 hours a week each, if you are doing it properly), and the manager hours spent checking output in the first few months.
Count them. A business case that ignores them looks better on the slide and falls apart in reality.
The reverse maths is also worth having ready. A finance person saving 2 hours a week pays back the licence several times over. Finance and operations tend to be where this returns fastest, which is also where I would start if you are choosing.
What Copilot is actually good at, and where it will fail your metric
This matters for measurement, because measuring the wrong process guarantees a bad number.
It is strong at analysis, investigation, summarising large messy inputs, and open questions. Give it 63 slides of budget control and ask what looks wrong, and it is genuinely useful.
It is not the right tool for a report that must come out identical on every run. Microsoft say this themselves. If your chosen process requires deterministic output, you picked the wrong process to measure and the number will punish you for it.
And every number that leaves the process gets checked by a person. It is the draft, you are the signature. Build that check into the process design, not as an afterthought when something goes out wrong.
Reviewing it without fooling yourself
3 questions at the 6 month review.
Did weekly usage hold, or did it peak and fade. Did the named process actually change, measured the same way as the baseline. And did anything change in how managers ask for work, because if managers are still asking for the same table they always asked for, the hours were spent and the benefit was handed back.
If usage is flat and nothing changed, resist the urge to buy more licences or run more training. Look at whether anyone in leadership actually required this, whether hours were allocated in writing, and whether it comes up in the management meeting. That is almost always where the answer is. I went through that failure pattern in why Copilot rollouts stall.
The 4 inputs that move the number
If the metric is disappointing, it is almost never the tool. It is one of these 4, and all of them sit with leadership.
Ownership. The CEO and the C-suite are the owners of this, not IT and not a committee. Committee-owned projects are the ones I watch end quietly around month 8.
Hours. Teams need time to learn, to keep learning and to practise, and champions need time to get good, to mentor and to teach others. In writing, not in spirit. If nobody can tell you how many hours were allocated and to whom, that is your answer.
A requirement, not encouragement. Leadership has to demand, genuinely demand, that people use it on every task or at least try. Using AI is not embarrassing any more, it is part of the job. And in the same breath: no copy and paste, because copy and paste is how you get publicly embarrassed. Every output is a draft, and what makes it worth sending is what you pour in, meaning insight, experience, creative thinking, strategy and a lot of human brain.
Presence in the routine. A standing item in management and team meetings. One recurring question, 3 minutes, works better than any company-wide email.
Fix those 4 and the numbers move. Skip them and you will be measuring the same flat line next quarter with better dashboards.
Questions I get about this
Can we measure ROI properly in the first quarter?
You can measure usage. You cannot honestly measure business impact yet. 3 to 6 months is realistic for a measurable change in how work gets done, and anyone promising you a P&L number in 6 weeks is selling something.
What if the numbers come back bad?
Then you learned it in month 6 rather than year 2, which is the whole point. In my experience a bad number is nearly always a management finding rather than a technology finding.
Should we benchmark against other companies?
Lightly. The 35.8% adoption average is useful for sanity. Everything else varies so much by sector and starting point that your own baseline is worth more than anyone else's case study.
Who should own the measurement?
Someone in the executive team, and not the same person who is selling the project internally. The separation matters more than the seniority.
Where I would start tomorrow
Write down the one process. Write down its current baseline, even roughly, even if it takes 20 minutes of asking people. Write down the date you will look again, and the 2 numbers.
That is a page of work, and it puts you ahead of the 86% who never define the impact at all.
If you want a hand structuring it for a specific organisation, that is part of what I do. There is more on the Microsoft Copilot consultant page.
Every live online format and price sits on one page: live online Copilot workshops and lectures.
Eyal Marcus is an AI consultant, trainer and keynote speaker who helps organisations actually adopt AI, working in English and Hebrew with enterprises in Israel and across Europe. He has been working with AI since early 2022, almost a year before ChatGPT launched, publishes the weekly "Don't Panic" AI newsletter and hosts the "Hands On AI" podcast. He has delivered 140 sessions in 55 organisations across healthcare, finance, industry, technology and the public sector.