- Record how a workflow performs before AI is introduced, using a method you can repeat later.
- Use leading measures such as adoption and task time early, and lagging measures such as unit cost later.
- Track hours, quality, cycle time, adoption and full cost together so speed is not bought with errors.
- At ninety days give each workflow one verdict: scale it, adjust it or stop it.
Most businesses that adopt AI cannot say, a few months in, whether it has paid for itself. They have licences, some enthusiastic users and a feeling that things are faster. When the finance team asks for evidence, the answers are anecdotes. That makes it hard to decide whether to expand, adjust or stop.
The fix is not complicated. It needs a baseline taken before anything changes, a small set of measures agreed in advance, and a clear decision point at the end of ninety days. This guide sets out how a CFO or finance lead can put that in place without building a reporting burden bigger than the work being measured.
Take a baseline before anything changes
You cannot measure improvement without knowing where you started. Before any AI tool is introduced into a workflow, record how that workflow performs today. How long does a typical task take? How many are completed each week? How often does work come back for correction? How long does a customer or colleague wait for the result?
This does not need to be perfect. A fortnight of simple timing by the people doing the work, or a sample pulled from your existing systems, is usually enough. What matters is that the same method is used again later, so you are comparing like with like.
Imagine a clinic group that wants to use AI to draft patient referral letters. Before starting, it records how long each letter takes, how many are sent back by doctors for changes, and how many days pass between consultation and letter. Three numbers, gathered once, give it something real to compare against.
Leading and lagging measures
Lagging measures tell you whether the investment worked: cost per transaction, revenue per employee, margin on a service line. They matter most, but they move slowly and are affected by many things besides AI.
Leading measures tell you early whether you are on track: how many people are using the tool each week, how many tasks are being run through it, how much time each task takes now. They move within weeks and help you correct course before the quarter is lost.
A sound measurement plan uses both. Leading measures guide the first ninety days. Lagging measures confirm, over the following quarters, whether the value is showing up in the accounts.
Five measures worth tracking
Keep the list short. These five cover most workflows.
- Hours: the time spent on the task before and after, measured the same way each time.
- Quality: the rate of errors, rework or returned items, so that speed is not bought at the cost of accuracy.
- Cycle time: how long the whole process takes from start to finish, including waiting, which is often where the real delay sits.
- Adoption: how many of the intended users are actually using the tool, and how often.
- Cost: the full cost of the AI work, including licences, set-up, training and the time people spend reviewing output.
If any one of these gets worse while the others improve, you have learned something important. Faster work with more errors is not a saving.
Turning hours into money, honestly
Time saved is only worth something if it is used. If a finance assistant saves several hours a week but those hours are simply absorbed, the business has not gained much. Be explicit about where the time goes: more invoices processed without new hires, faster month-end close, more sales calls made, or overtime reduced.
Value the hours at the loaded cost of the people involved, not just their salary, and count only the hours that have been redeployed to something useful. This produces a more modest number than a vendor's calculator, but it is one the board will trust.
A saving that cannot be pointed to in the accounts is a hope, not a return.
A simple ninety-day rhythm
In the first month, take the baseline, set up the tool on one or two workflows, and train the people involved. Track adoption weekly. Expect slower work for a short period while people learn.
In the second month, measure hours, quality and cycle time against the baseline. Look for workflows where adoption is low and find out why. Often the cause is a poor fit with how the work is actually done, not resistance.
In the third month, compare the full cost against the value of redeployed time and quality gains. Prepare a short summary for the leadership team, written in plain language with the evidence attached.
Deciding to scale, adjust or stop
At the end of ninety days, every workflow should get one of three verdicts. Scale it where the measures show clear gains and people are using it willingly. Adjust it where results are mixed, by changing the process, the prompts or the training, and set a short further review. Stop it where the gains are not there, and record why, so the lesson is not lost.
Stopping is not failure. A disciplined decision to stop a weak use case frees budget and attention for a stronger one, and it shows the team that AI investment is judged on evidence like any other spend.
Getting the first measurement right
The hardest part is usually choosing which workflows to measure first, and setting a baseline people trust. If you would like help with that, our complimentary one-week Cost Review identifies the workflows where AI is most likely to pay back, sets the baselines, and gives you a ninety-day plan with the measures already defined, whether or not you choose to work with us.
Want these insights applied to your business?
Our complimentary one-week Review finds your biggest leaks and gives you a 90-day plan, whether or not you work with us.





