Usage is not quality: what to measure when you roll out AI
Adoption figures show who clicked. They do not show whether the work got better, or what it cost to check.
When an organisation rolls out an AI tool, the first numbers to arrive are usage numbers. Active users. Prompts per week. Teams onboarded. They are easy to collect, they tend to go up, and they fit neatly on a slide.
They answer a real question: are people trying the tool? They do not answer the question that matters: is the work better, and at what total cost? Those are different questions, and they need different measures.
What usage can and cannot tell you
Usage measures behaviour, not outcomes. High usage might mean the tool is valuable. It might also mean novelty, encouragement from above, or people using it for tasks where it is weak. Low usage might mean the tool is poor. It might also mean the tool sits in the wrong place in the workflow. Or the people who would gain most have not been shown how to use it.
You cannot tell which from the count alone. So treat usage as a measure of exposure, not success. It tells you how many outputs are being produced, and so how many need to be right. Rising usage with no view of quality is rising risk, not rising value.
Define good before you measure
Start with the task, not the tool. Before you look at any AI output, write down what a good result looks like for this piece of work. Use terms someone else could check.
For a summary of evidence, that might be three things. Faithfulness: does every statement match what the sources say? Coverage: does it include the points a careful reader would expect? Misstatements: does it assert anything the sources do not support? For a drafted letter, it might be accurate facts, the right tone and everything the letter is required to contain.
Then build a small test set that looks like the real work. Use realistic inputs, with reference answers or marking criteria agreed by people who know the task. Keep it fixed, so you can run it again when the model, the prompt or the workflow changes. Some criteria can be checked by code, such as whether every citation points to a source that exists. Others need judgement, from a person or from a model grader whose agreement with people you have measured. Say which is which.
The OWASP Top 10 for LLM Applications lists misinformation as a risk in its own right. In the current 2026 edition it is LLM07, and its examples include misleading summaries and critical omissions. The entry also names overreliance as a key factor: people often treat fluent, confident output as authoritative. Measuring quality directly is how you find out whether that risk is real in your setting, rather than assuming it away.
Count all the effort
Speed of generation is not time saved. A draft that appears in seconds but takes an hour to verify may cost more than writing it by hand.
Measure the total human effort for each output that is accepted. That includes writing the prompt, reading the draft, checking it against sources, correcting it, and any rework when an error is found later. Record checking time as part of the task, not as overhead. Record it as people work, too. Effort reconstructed from memory at the end of the week is unreliable.
Two traps are common. The first is timing only the generation step, which makes every tool look fast. The second is treating skipped checks as savings. If people stop checking because the output looks fluent, you have not saved effort. You have moved it into risk that nobody is measuring.
Baselines are not test results
Two kinds of number often get mixed up.
A baseline describes how the work performs today, without the tool. It covers how long the work takes, how often it needs correcting, and what errors reach the people who rely on it. Without a baseline, “better” has nothing to be better than.
A test result describes how the AI-assisted version performs on a defined set of tasks, under conditions you control. It is useful and repeatable. It is not the same as performance in live use, where inputs are messier and people behave differently.
Be clear about which one you have. A good test result on a curated set does not prove a benefit in live use. A self-reported time saving from a pilot group, with no baseline to compare against, proves very little. Neither is worthless. Both need to be labelled for what they are.
Report small samples honestly
Early evaluations are small. That is fine. Pretending otherwise is not.
Report counts, not just percentages. Writing “seven of ten summaries had no misstatements” tells the reader how much weight the result can bear. Writing “70%” hides it. Show the failures, with examples, because they are often more informative than the average. If you give a rate, give a sense of how uncertain it is. With a handful of cases the uncertainty will be wide. Saying so is rigour, not weakness.
Say what you did not test: task types, user groups, document lengths, languages. A reader should be able to see where the evidence stops.
Match the effort to the stakes. A low-risk drafting aid does not need a trial. A tool whose output feeds into advice or decisions needs more than a satisfaction survey. When you need more than a local comparison, start with HM Treasury’s Magenta Book, the central government guidance on evaluation.
A short checklist
Before you report on an AI roll-out, check that you can answer these questions.
- What does a good output look like for this task, and who agreed that?
- What is the baseline without the tool?
- How good are the outputs against a fixed test set, and what failed?
- What is the total human effort for each accepted output, including checking and correction?
- How many cases is each number based on, and what was not tested?
- Is the depth of evaluation proportionate to the risk?
If you can answer these, usage figures become useful context. If you cannot, usage figures are the only story you have, and it is not the one you need.
Sources
- OWASP GenAI Security Project, LLM07:2026 Misinformation, OWASP Top 10 for LLM Applications 2026 (source text on GitHub). In the 2025 edition this entry was LLM09:2025 Misinformation.
- HM Treasury, The Magenta Book, GOV.UK. Central government guidance on evaluation.
Views are my own. Where I refer to public guidance, this is my reading of it: unofficial; not government guidance.