Insights · AI Consulting
How to measure an AI consulting engagement.
Most AI projects get judged on whether they shipped. The better test is whether a number moved. The measures an Australian organisation should agree before an AI consulting engagement starts, and how to read them afterwards.
Ask an organisation how its last AI project went and you usually hear about the launch: the demo worked, the pilot went live, the board saw it. Ask what changed in the work six months later and the answer gets vaguer. That is rarely because nothing changed. It is because nobody agreed, before the engagement started, what change would count.
Measuring an AI consulting engagement is not hard. It just has to be done in the right order: pick the measure, take the baseline, then build. Most of the disappointment we see in AI projects comes from doing those three things the other way round. This piece sets out what to measure, when to read it, and what to write into the contract so the measure survives the engagement.
Agree the measure before the work starts.
Once a project is under way, every team finds a number that makes its own work look good. That is human, not dishonest. The fix is to settle the measure while nobody has anything to defend yet, ideally in the brief itself. Our guide on how to write an AI consulting brief puts it second on the list for that reason: name the decision or workflow that changes, and you have named what to measure.
A useful measure has three properties. It describes the work, not the technology. You can count it today, before anything is built. And the people who do the work would agree that moving it is worth paying for. "Time to reconcile a month of claims" passes all three. "Number of AI use cases identified" passes none.
Three layers of measure.
Most engagements need one measure at each of three layers. Each answers a different question, and they move at different speeds.
- The work measure. What happens to the task itself: time per case, cases handled per week, the share that need rework, the share that breach a rule. This moves first, and it is the one to watch closely.
- The decision measure. Whether people decide better or sooner: how long a rostering gap stays open, how many days before month end the forecast is trusted, how often a manager overrides the recommendation. This moves second.
- The money measure. What it is worth: hours released and where they went, revenue recovered, write-offs avoided, cost per completed task. This moves last, and it is the one most often claimed too early.
If an engagement can only name a money measure, be careful. A money figure with no work measure underneath it is a forecast, not a result.
Take the baseline first.
A measure without a baseline is a number with nothing to compare it to. Before anything is built, record two to four weeks of the current process: how long the work takes, how often it goes wrong, and what it costs. Use the same definitions you will use afterwards, and write them down.
Sometimes the baseline cannot be taken, because the data to count the work does not exist or the systems disagree about it. That is not a delay to the project. It is the first finding, and often the most valuable one, because an AI system will run into the same gap. An AI readiness assessment is largely the work of finding that out early, before you pay for a build that depends on records you cannot trust.
Numbers that look good and mean little.
Some figures turn up in nearly every AI project report. None of them is wrong, but none of them tells you whether the engagement worked.
- Model accuracy on a test set. Useful to the people building it. It says nothing about whether the output is right on your live cases, or whether anyone acts on it.
- Hours saved, with no destination. If the released hours are not visibly spent on something else, they were not saved. Ask where they went.
- Licences activated or users onboarded. A measure of rollout, not of value.
- Use cases identified. A list is an input to a decision, not an outcome.
- A benefit range with no owner. A figure that nobody on your side has agreed to be measured against is marketing, whoever wrote it.
What to measure for an AI agent.
An AI agent does work rather than advising on it, so it needs measures closer to those you would use for a new team member than for a report.
- Completion without correction. The share of tasks the agent finishes that a person accepts as they are.
- Correction and override rate. How often a person changes or rejects what it did, and why. The reasons matter more than the rate.
- Time to approve. If every output needs a long review, the work has moved, not shrunk.
- Cost per completed task. Licence, usage, upkeep and review time together, divided by tasks done. Our piece on agent platform or AI consulting firm lists the costs a licence line leaves out.
- Drift. The same measures, re-read after a rule, price or policy changes. An agent that was right in March can be wrong in July without anyone touching it.
For a care provider checking claims, that might mean the share of flagged claims that really were wrong, the days taken to reconcile a month, and how often a person overrides the check. Our pieces on AI for NDIS providers and AI for home care providers describe the work those measures would sit on.
When to read the numbers.
Agree the reading points at the start, and agree what happens at each one.
- After about 30 days: is it used? If the people meant to use it have gone back to the old way, find out why before reading anything else.
- After about 90 days: has the work measure moved? For a well-scoped first project, it should have. If it has not, stop and find out why before spending more.
- After six to twelve months: what is it worth? Only now is the money measure worth reading, because the work has settled and the seasonal noise has had a chance to wash out.
The stop rule matters as much as the targets. An engagement that cannot fail on a measure cannot succeed on one either. Our guide on when to bring in an AI consultant makes the same point from the other end: decide in advance how you will know it worked.
Write it into the statement of work.
A measure that lives only in a kick-off slide does not survive the first change of scope. Put four things in the contract.
- The measures, with definitions. Exactly what is counted, from which system, and by whom.
- The baseline, or the work to take it. If it does not exist yet, taking it is the first deliverable.
- The reading points and the stop rule. When the numbers are read, and what result pauses the work.
- Who owns the measure afterwards. A named person on your side, with the dashboard or report that keeps it running once the consultants leave.
What an AI consulting engagement looks like covers the rest of what a statement of work should say, and what AI consulting actually delivers sets out what you should own when it ends.
Why RUBIX.
RUBIX is an independent Australian data and AI consultancy, founded in 2011. That is 15 years, 450+ projects and 115+ customers, and much of that work has been the unglamorous part of measurement: getting an organisation's records to agree with each other so a baseline can be taken at all. Three things follow from that for how we run an engagement.
- Data first, so the measure is real. We build the data foundation that joins your systems, so the work measure comes from one governed source rather than a spreadsheet assembled for the steering committee.
- We say what kind of number it is. For a leading NDIS provider we joined six disconnected systems into one governed data foundation, a live role-based dashboard and an automation roadmap that quantified around $645K of annual impact, inside eight weeks. That figure is an estimate the roadmap was built to test, not a result, and we describe it that way. Read the NDIS provider case study.
- Senior people stay with the numbers. Our forward deployed engineers sit with your team, so the people who agree the measure are the people who deliver against it. That is how our AI consulting in Australia is set up.
If you are planning an AI project, start with a conversation about what number should move and whether you can count it today. Sometimes that conversation is the most useful part.
Frequently asked questions.
How do you measure whether an AI consulting engagement worked?
Agree one or two measures of the work itself before the engagement starts, such as time per case or the share of cases that need correcting, take a baseline of those measures from your current process, and compare against that baseline at fixed points afterwards. If nothing was agreed up front, any result can be described as a success.
What should you measure for an AI agent?
Measure how often it completes a task without a person correcting it, how often a person has to correct or override it, how long approvals take, and what each completed task costs once usage, upkeep and review time are counted. Keep measuring after go-live, because rules and data change.
How long before an AI project shows a return?
Usage should be visible within the first month, and the work measure should move within about three months for a well-scoped first project. The money measure usually takes six to twelve months to read honestly. If the work measure has not moved by three months, stop and find out why before spending more.
General information only, not legal, financial or regulatory advice. Current as at October 2026.