The decision a baseline record actually settles
Before an AI tool is used in a process for the first time, there is exactly one way to later prove that something genuinely changed: record the current state before deployment happens. What is at stake is whether, after a month of use, there is a comparable figure to point to, rather than just an impression that work feels faster or more comfortable. The process owner, the measurement period and the exact definition of a unit of work need to be agreed in advance, because the baseline cannot be reconstructed after the fact.
A staged rollout of a generative assistant among more than five thousand customer support workers, studied by Brynjolfsson, Li and Raymond, is a rare case where a before and after comparison was actually documented in detail, with a reported average rise in issues resolved per hour and considerable variation by worker experience. It is one company’s one process, not a promise of what your own baseline will show, and the point here is narrower: without a comparable record of your own starting point, that kind of comparison is not available to you at all.
The unit of work and what to record for it
The first step is naming the unit of work concretely enough to count: one processed invoice, one reply to a customer enquiry, one contract draft, one meeting note. A vague unit such as “work on the project” compares badly, because every project has a different scope.
Four measures get recorded for each unit:
- Full time. From taking on the assignment to output ready for standard review, including time spent waiting for input, approval, or a colleague’s reply, and including any rework. Active work time alone understates the real length of the process, and waiting time is often the first thing that disappears after an AI tool is deployed.
- Quality against a consistent measure. Use an existing acceptance criterion if one exists. If there is no consistent criterion, build a simple three-point scale, for example no changes needed, minor edit, returned for rework.
- Number of corrections. How many times the output was sent back before being accepted as finished, and who made the correction.
- Serious failures. Cases where the output contained a factual error that routine acceptance without a thorough check would likely have missed, or where the process failed outright and had to restart from the beginning.
How many samples and over what period
There is no universal sample size, because it depends on how often the unit of work repeats and how much its difficulty varies. For a task a team handles daily, recording cases across two to four weeks of normal operation is usually realistic. For a task that only recurs a few times a month, plan for a longer period, or the sample will only capture one type of situation.
The chosen period should cover the normal range of difficulty, meaning a mix of straightforward and demanding cases from everyday operation. If the team only occasionally deals with seasonal peaks or unusually hard cases, note that alongside the data so a later comparison can account for the context.
A short period with a small sample works well as a first check on whether the record can even be kept in practice. A decision to continue or expand needs a longer, more representative collection of data.
What has to stay unchanged for the comparison to mean anything
A comparison between the baseline and the post-deployment state only means something if nothing else material changes between the two measurements. This applies in particular to:
- The definition of the unit of work and the quality criteria. If what counts as a finished output changes in the meantime, or acceptance gets stricter, the figures stop being comparable.
- The pool of people performing the task. Swapping an experienced team member for a newcomer changes the result regardless of the tool.
- The input material and its quality. If the data source, form, or system the team draws from changes in between, the comparison picks up that change too.
- The scope of review. If an approval step is added or removed after the AI tool is deployed, that change shows up in the measured difference alongside the effect of the tool itself.
Tracking spend across the tools involved in a deployment like this is a related but separate exercise, covered in how to manage a budget across multiple AI tools, and the wider cost picture a baseline record eventually feeds into is covered in what AI deployment really costs a company.
Card for recording the baseline
The sheet below is for a single, consistent record. It gets filled in separately for each recorded unit of work, which later allows you to calculate the spread of results as well as the average.
| Field | What to record |
|---|---|
| Date and unit of work | Specific identification of the unit being processed, for example an invoice number or an enquiry reference. |
| Who performed the task | Role or name, per your internal agreement on record anonymisation. |
| Start and handover time | Exact time the assignment was taken on, and the time the output was handed over for review. |
| Full time in minutes | Sum of active work, waiting, and any rework. |
| Quality rating | Rating on the agreed scale, for example no changes needed, minor edit, returned for rework. |
| Number of corrections | How many times the output was sent back before being accepted as finished. |
| Serious failure (yes/no) | Whether the output contained a factual error, or the process had to restart from scratch. |
| Note on circumstances | Anything unusual about this case, for example a missing input or an unusual deadline. |
The row below is a made-up illustrative example of how to fill in the sheet, not actual measured data:
| Field | Example value (illustrative, not real data) |
|---|---|
| Date and unit of work | 3 March, invoice no. 118 |
| Who performed the task | accounts clerk A |
| Start and handover time | 9:14 to 10:02 |
| Full time in minutes | 48 (12 minutes of that spent waiting on supplier confirmation) |
| Quality rating | minor edit |
| Number of corrections | 1 |
| Serious failure | no |
| Note on circumstances | supplier sent the invoice in the wrong format, line items had to be retyped by hand |
What a before and after comparison proves, and what it does not
The difference between the baseline and a later measurement shows that something in the process changed, and by how much. On its own it does not establish the cause. Other things could have changed in the same period: people getting better trained, an input form being revised, a seasonal dip in work volume, or simply a team getting faster at a task through routine, independent of any tool.
Recording what else changed in the meantime raises the value of the comparison. So does comparing across several people or a longer period rather than a single before and single after figure. Even then, this remains an observation from one process at one company. It is a useful basis for an internal decision to continue, adjust, or stop a pilot. On its own it is not general proof that the same tool will produce the same effect anywhere else. A full, structured review of an entire AI deployment, rather than one process’s baseline, is a separate and broader exercise, covered in what an AI system audit covers and how it works.
Where this record stops and training-impact measurement begins
This method concerns the baseline of a process before an AI tool joins the work. It does not concern how much people’s own knowledge has changed after training. If a company is primarily training staff and wants to evaluate the teaching itself, a different methodology and different criteria apply, closer to what is covered in how often to repeat AI training and when a refresher is enough. The two measurements can meet in practice if an AI tool is rolled out alongside training, but the record described here serves the process and its output. It does not cover evaluating the training itself.
Sources and limits
The NIST AI Risk Management Framework was verified on 5 August 2026. By NIST’s own account, it continues to be revised. The cited Measure function and TEVV testing activities describe a voluntary vocabulary for risk management. This is not a mandatory procedure and not a certification standard. The Brynjolfsson, Li and Raymond study documents only the outcome of one company’s one deployment of a generative assistant in customer support. It does not represent an expected outcome for a different process or a different company, and this article does not present it as one. The method described here is a CIAD recommendation for a company’s own internal record. It is not a certified methodology and not a legal requirement to measure anything at all.
Frequently asked questions
How many samples do I need for a baseline measurement before AI?
The right number depends on how often the task repeats and how much its difficulty varies. For a task a team handles daily, recording cases across two to four weeks of normal operation is usually manageable. The goal is not statistical certainty. You need enough cases that one unusually easy or hard day does not distort the whole picture.
Do I need to measure waiting time, not just the active work itself?
Yes. If you only record active work time, you lose exactly the part of the process an AI tool most often changes: waiting for input, approval, or a colleague's reply. The full time from assignment to output ready for review is the one figure you can honestly compare later against the post-deployment state.
What if something else changes during the baseline measurement, such as team staffing?
The comparison loses value, because you will not be able to tell whether a later difference came from the AI tool or from a change in people, inputs, or deadlines. It is best to choose a measurement period where the process, the team and the quality criteria stay unchanged. If that is not possible, record the change alongside the numbers and account for it in the comparison.
Is one experienced person enough to run the baseline measurement?
One experienced person makes the measurement easier to run, but the result then describes only their pace and working style. If more than one person performs the task, the sample should include at least some of them. Otherwise the comparison after AI deployment will show more of a difference between individual people than between the two ways of working.
How do I record quality without a formal rating scale?
Build a simple scale of your own from existing acceptance criteria, for example no changes needed, minor edit, returned for rework. What matters is using the same scale before and after, and having the same role that already judges the output continue to judge it. Changing the reviewer mid-measurement would distort the result just as much as changing the criteria itself.
SOURCES AND VERIFICATION
reviewed Lukáš Dlouhý ·
- NIST AI Risk Management Framework 1.0 (verified 5 August 2026)
- NIST AI RMF Core: Govern, Map, Measure, Manage (verified 5 August 2026)
- NIST on test, evaluation, validation and verification of AI (verified 5 August 2026)
- Brynjolfsson, Li and Raymond, Generative AI at Work, Quarterly Journal of Economics 140(2), 2025 (verified 5 August 2026)