# AI Workflow Baseline and After Measurement Card

Use this worksheet to compare one recurring workflow before and after an AI-assisted change. It is designed for a small pilot where the team needs enough evidence to make a decision without turning saved minutes into invented ROI.

Copy the blank template into your working document or repository. Agree on the measures and decision rules before the first pilot example. Keep observed results separate from projections.

## How to use the card

1. Pick one workflow with a recognizable trigger, output, and reviewer.
2. Record three to five recent baseline examples. Use more when the task varies widely or the decision carries more risk.
3. Define the clock, quality bar, review measure, and safety checks before testing the new path.
4. Run comparable examples through the pilot. Do not discard slow cases or failures.
5. Compare medians and rates. Show the sample size beside every result.
6. Ask the person doing the work and the reviewer what became easier or harder.
7. Make the pre-agreed decision: scale, revise, contain, or stop.

A small sample can guide the next test. It does not establish population-wide productivity, booked savings, or a causal business result.

## Blank reusable template

### 1. Pilot definition

| Field | Entry |
|---|---|
| Workflow |  |
| Trigger |  |
| Output that moves to the next real step |  |
| Workflow owner |  |
| Reviewer or decision owner |  |
| Baseline date range |  |
| Pilot date range |  |
| Baseline sample target |  |
| Pilot sample target |  |
| Tool, model, and version if known |  |
| Reusable artifact being tested |  |
| Approved input type or data classification |  |
| What this pilot will not measure |  |

### 2. Measurement rules

Write these rules before collecting results so the team cannot move the finish line after seeing the data.

**Time to usable result**

- Clock starts when:
- Clock stops when:
- Unit: minutes / hours / days
- Summary: median, with minimum and maximum shown
- Pauses or waiting time included:

**Quality bar**

Score each example as `Pass`, `Needs edits`, or `Fail` against three to five criteria.

| Criterion | Pass means | Needs edits means | Fail means |
|---|---|---|---|
| 1. |  |  |  |
| 2. |  |  |  |
| 3. |  |  |  |
| 4. |  |  |  |
| 5. |  |  |  |

**Review effort**

- Measure: reviewer minutes / correction rounds / number of substantive fixes
- What counts as a substantive fix:
- Who records it:

**Safety and policy checks**

Mark every required check `Pass` or `Miss` for every example. A miss stays visible even when averages improve.

| Required check | Evidence to retain | Is one miss a stop condition? |
|---|---|---|
| Approved input and account used |  | Yes / No |
| Factual claims trace to a source |  | Yes / No |
| Required disclosure present |  | Yes / No |
| Human approval completed before external action |  | Yes / No |
| Other: |  | Yes / No |

**Qualitative question**

- Person doing the work: What would stop you from using this on the next real example?
- Reviewer: What became easier or harder to verify?

### 3. Per-example log

Keep baseline and pilot examples in the same table. Give each example an anonymous ID rather than pasting sensitive inputs into the scorecard.

| Phase | Example ID | Date | Time to usable result | Quality | Review effort | Safety | Exception or observation |
|---|---|---|---:|---|---:|---|---|
| Baseline |  |  |  |  |  |  |  |
| Baseline |  |  |  |  |  |  |  |
| Baseline |  |  |  |  |  |  |  |
| Baseline |  |  |  |  |  |  |  |
| Baseline |  |  |  |  |  |  |  |
| Pilot |  |  |  |  |  |  |  |
| Pilot |  |  |  |  |  |  |  |
| Pilot |  |  |  |  |  |  |  |
| Pilot |  |  |  |  |  |  |  |
| Pilot |  |  |  |  |  |  |  |

### 4. Side-by-side summary

| Measure | Baseline | Pilot | Observed difference | Important limit |
|---|---:|---:|---:|---|
| Sample completed |  |  |  |  |
| Median time to usable result |  |  |  |  |
| Time range |  |  |  |  |
| Quality passes |  |  |  |  |
| Quality needs edits |  |  |  |  |
| Quality failures |  |  |  |  |
| Median review effort |  |  |  |  |
| Safety checks passed |  |  |  |  |
| Safety misses |  |  |  |  |

### 5. Decision rules agreed before the pilot

Replace the blanks with thresholds that fit the workflow.

- **Scale:** Median time improves by at least ___, quality pass rate is at least ___, review effort does not increase, and every stop-condition safety check passes.
- **Revise:** The result improves on ___, but a named quality, review, access, or safety issue needs another test.
- **Contain:** The method works only for ___ role, input type, or risk level. Keep it there and review again on ___.
- **Stop:** The result does not improve enough to justify ___, or the pilot crosses this stop condition: ___.

### 6. Decision record

| Field | Entry |
|---|---|
| Decision: Scale / Revise / Contain / Stop |  |
| Evidence supporting the decision |  |
| Exceptions and failures |  |
| What remains unknown |  |
| Owner of the next action |  |
| Next review date |  |

### 7. Evidence to attach

- [ ] Completed per-example log with sample size
- [ ] One representative baseline output
- [ ] One representative pilot output
- [ ] One failed or exception example
- [ ] Quality rubric used by the reviewer
- [ ] Reusable Skill, checklist, or workflow card tested
- [ ] Safety or approval receipt where required
- [ ] Notes from the worker and reviewer questions

## 25-minute lab exercise

Use this exercise with sanitized or synthetic inputs.

1. **Five minutes:** Read the scenario and define when the clock stops. If the group cannot agree on "usable," fix that before measuring.
2. **Seven minutes:** Choose three quality criteria, one review measure, and the safety checks that can block scaling.
3. **Eight minutes:** Read the results and make one decision using rules written before the reveal.
4. **Five minutes:** Compare the group's decision with the completed synthetic example below. Name the evidence that changed the decision.

The exercise is complete when the group can explain why it chose scale, revise, contain, or stop. Finishing the table is not the outcome.

## Completed worked example

**This example is fully synthetic. The team, workflow, inputs, and results are fictional. It demonstrates the worksheet and is not client evidence.**

### Fictional scenario

An operations team turns sanitized weekly meeting notes into an internal project recap. A coordinator drafts the recap, then a project lead reviews it before posting it to the internal workspace. The pilot tests a reviewed Skill that asks for a fixed structure, labels unsupported details, and requires an owner and date for every action.

The team records five baseline examples and five pilot examples. The small sample is enough to decide whether to revise and test again. It cannot support a company-wide productivity claim.

### Rules agreed before the test

- Clock starts when the coordinator opens the approved notes.
- Clock stops when the project lead says the recap is ready to post.
- Quality is `Pass` only when the recap has the required sections, every factual statement traces to the notes, status language is neutral, and every action has an owner and date.
- Review effort is the project lead's active review time in minutes.
- Safety passes only when the approved notes and account are used, every claim traces to the notes, and the lead approves before posting.
- Scale requires at least 20% lower median time, at least 80% quality passes, no increase in median review time, and 100% safety passes.
- Revise applies when time and quality meet the bar but a fixable safety or review issue remains.

### Synthetic example log

| Phase | Example ID | Time to usable result | Quality | Review minutes | Safety | Observation |
|---|---|---:|---|---:|---|---|
| Baseline | B-01 | 62 min | Pass | 18 | Pass | No exception |
| Baseline | B-02 | 55 min | Needs edits | 16 | Pass | One action lacked an owner |
| Baseline | B-03 | 71 min | Pass | 22 | Pass | Long notes |
| Baseline | B-04 | 58 min | Needs edits | 17 | Pass | Status language overstated progress |
| Baseline | B-05 | 64 min | Pass | 19 | Pass | No exception |
| Pilot | P-01 | 41 min | Pass | 12 | Pass | No exception |
| Pilot | P-02 | 38 min | Pass | 10 | Pass | No exception |
| Pilot | P-03 | 47 min | Needs edits | 15 | Miss | Skill inferred a due date absent from notes |
| Pilot | P-04 | 43 min | Pass | 13 | Pass | Long notes |
| Pilot | P-05 | 39 min | Pass | 11 | Pass | No exception |

### Synthetic side-by-side result

| Measure | Baseline | Pilot | Observed difference |
|---|---:|---:|---:|
| Sample completed | 5 | 5 | Same sample count |
| Median time to usable result | 62 min | 41 min | 21 min lower, about 34% |
| Time range | 55 to 71 min | 38 to 47 min | Narrower in this sample |
| Quality passes | 3 of 5 | 4 of 5 | 60% to 80% |
| Median review effort | 18 min | 12 min | 6 min lower |
| Safety checks passed | 5 of 5 | 4 of 5 | One pilot miss |

### Synthetic decision: Revise

The pilot met the pre-agreed time, quality, and review thresholds. It did not meet the 100% safety condition because one recap invented a due date that was not in the approved notes.

The next version should require missing dates to appear as `Date needed` and add a deterministic check that rejects any dated action without a source line. The team should rerun five comparable examples, including one where dates are missing. It should not scale the workflow until every stop-condition safety check passes.

That decision is less dramatic than declaring a 34% productivity gain. It is also more useful. The observed speed difference justifies another test, and the safety miss says exactly what has to change first.

## Keep the claim smaller than the evidence

Report what happened in the sample, what did not, and which decision followed. A defensible statement looks like this:

> In five synthetic pilot examples, median time to a reviewed recap was 21 minutes lower than in five synthetic baseline examples. Quality met the pre-agreed bar, review time decreased, and one safety check failed. The decision was to revise and test again.

If you project the result to a larger volume, put the projection in a separate box with its assumptions. Returned time is not automatically booked savings, and a workshop result is not proof that a workflow survived normal work.

Related reading: [How to Measure AI Adoption Without Inventing ROI](/blog/measure-ai-adoption-without-fake-roi) and [How to Run Safe AI Labs on Real Team Work](/blog/safe-ai-labs-real-team-work).
