The Dashboard Isn't the Outcome
AI adoption dashboards are good at counting visible activity: active users, chats, generated lines, accepted suggestions, minutes in the tool.
Those numbers answer a procurement question. They tell you whether a license has been touched. They don't tell you whether a client brief got better, a review got shorter, a decision got safer, or a team learned a workflow it can repeat.
That gap creates two bad stories.
In the first, a company sees high activity and declares success even though people are producing more work for reviewers to fix. In the second, a company can't attach a clean dollar figure to a four-week pilot, so it calls the pilot unmeasurable.
You don't need a made-up ROI model. You need a baseline, an observable change, and a decision the evidence can support.
The unit of measurement should be the workflow, not the tool.
Write the Before Card First
Pick one recurring task. Describe how it happens before the new workflow starts.
A useful before card fits on one page:
- Trigger: What starts the work?
- Inputs: What files, notes, systems, or decisions does it require?
- Current path: What are the main steps?
- Time: How long does a normal example take?
- Quality: What makes the result usable?
- Review: Who checks it, and what do they usually fix?
- Risk: What must never be skipped?
- Volume: How often does the team do this?
Don't turn the baseline into a research program. Three to five recent examples are usually enough to see the shape. Use a median when one case is unusually easy or painful. Ask the person doing the work and the person reviewing it. Their answers often expose different costs.
Write uncertainty directly on the card. "Usually 45 to 75 minutes" is more honest than "60 minutes" if nobody has timed it. "Reviewer rewrites the recommendation section in most cases" is useful even before you have a percentage.
The before card gives the team something concrete to beat. It also stops the pilot from quietly changing the definition of success after the results arrive.
Measure More Than Speed
Time saved is attractive because it sounds precise. On its own, it can reward worse work.
Track a small set of dimensions that match the task:
Time to a usable result. Start the clock when the work begins and stop when the output can move to its next real step. Don't stop at the model's first answer if a human then spends 20 minutes fixing it.
Quality against a shared bar. Use three to five criteria the team already cares about: completeness, accuracy, tone, evidence, or decision usefulness. A simple pass, needs edits, or fail scale is often enough.
Review effort. Count correction rounds or estimate reviewer minutes. This catches workflows that feel fast to the creator because they push labor downstream.
Consistency. Check whether different people using the same method produce results that meet the same bar. A workflow that only works for the person who designed it isn't ready to spread.
Safety. Record whether required sources, disclosures, approvals, and data rules were followed. A safety miss is not a rounding error in the average score.
You rarely need every metric. Choose two primary measures and one guardrail. For a client recap, that might be time to usable draft, reviewer corrections, and no unsupported claims. For a research synthesis, it might be evidence coverage, reviewer confidence, and source traceability.
Separate Observed Value From Projected Value
Observed value is what happened in the pilot. Projected value is what might happen if the workflow scales.
Keep them in separate boxes.
An observed statement sounds like this: "Across six recaps, median time to a reviewed draft fell from 58 minutes to 37 minutes. The reviewer requested the same number of factual corrections and fewer tone edits."
A projection sounds like this: "At 80 recaps per month, the observed difference could return about 28 hours, assuming volume, review quality, and team behavior stay similar."
The second statement is useful because its assumptions are visible. It is not booked savings. It does not claim that every returned hour becomes revenue. It gives a leader enough information to decide whether a larger test is worth running.
Avoid multiplying a best-case time estimate by every licensed user and calling the result annual ROI. That number usually ignores adoption rate, task frequency, review time, and the fact that saved minutes don't automatically leave the cost base.
Good measurement makes the next bet clearer. It doesn't need to make the current bet look bigger.
Use a Side-by-Side After Card
At the end of the pilot, duplicate the before card and add the new path.
The after card should show:
- The workflow steps that changed
- The reusable artifact involved, such as a Skill or reference file
- The human review point
- Results from the same measures used in the baseline
- Exceptions and failures, not just the clean examples
- What the team learned about inputs, context, and tool choice
Put the two cards next to each other. A leader should be able to understand the change without sitting through a demo.
The artifact matters because it explains why the result might repeat. If the improvement came from one expert writing a perfect private prompt, you measured a person. If the team used a reviewed Skill with named inputs and a quality check, you measured the beginning of a shared capability.
This is where adoption and workflow performance meet. The team isn't just using the tool more. It has a method another person can pick up.
Add One Qualitative Question
Numbers miss friction that people can name immediately.
After each test, ask: "What would stop you from using this on the next real example?"
The answer may be access, trust, missing context, an awkward copy-paste step, unclear policy, or a result that still takes too much editing. Those answers are design inputs. Group them and count how often they appear.
Also ask the reviewer: "What became easier or harder to verify?" A workflow can reduce drafting time while making the evidence trail worse. The reviewer sees that before the dashboard does.
Qualitative evidence doesn't need to become a fake score. Keep the language, tag the pattern, and connect it to a change in the next version.
Make the Decision Before the Readout
Every pilot should end with one of four decisions:
- Scale it: The result improved and the guardrails held.
- Revise it: The workflow shows promise, but a specific failure needs another test.
- Contain it: The method works for a narrow role or task but shouldn't spread yet.
- Stop it: The value didn't justify the friction or risk.
Define those choices before the readout. For example: scale if median time falls by at least 20 percent, quality does not decline, and every safety check passes. Revise if time improves but review effort increases. Stop if unsupported claims appear in more than one test.
The thresholds don't need to be universal. They need to be agreed, relevant, and visible.
That protects the team from two forms of theater: a success story built from the best example, and a permanent pilot that never earns a decision.
A Small Scorecard Beats a Big Claim
For most business-team pilots, the whole scorecard can be six lines:
- Workflow and owner
- Baseline time to usable result
- New time to usable result
- Quality or review difference
- Safety result
- Decision and next review date
Attach two example outputs, including one that failed. Link the reusable artifact. State the sample size. Note what you couldn't measure.
If adoption is uneven before the pilot starts, diagnose that first in Your Team Has AI Access. Why Isn't Anyone Using It?. If the next step is a controlled practice session, How to Run Safe AI Labs on Real Team Work lays out the room design.
The honest version of AI ROI is usually less dramatic than the sales version. It is also much more useful: this workflow got faster, the quality held, the review path is clear, and here is the evidence for the next decision.