Skip to content
OperateAgentic

HomeGuidesBuild AI

Guide · Build AI

What to measure without fake ROI

Track time on the task, the errors you can point to, and whether operators trust the draft. Leave invented percentages off the page.

3 min read

Spectrum Warm guide frame with a butter underline and a four-step path.

Small businesses get sold numbers they did not collect. A pilot does not need a fantasy lift to be worth keeping. It needs evidence the people doing the work recognize.

In short. Measure the task, the mistakes you can show, and operator trust. Do not publish a percentage you calculated backward from a hope.

Time on the task

Time is useful when it is the same task, before and after, on real instances. “Quoting a standard job” before the pilot and during the shadow. Not “the business feels faster.”

Write minutes, or a range, from a handful of instances. If the range is wide, say so. Wide ranges are what small shops actually have. A single average hides the job that took all afternoon because the customer changed scope.

Include the checkpoint in the time. A draft that is quick and a review that is long is not a win until you count both. Many tools look fast because they moved the minutes onto someone else.

Errors you can point at

An error is a specific miss: a wrong part, a promise the shop cannot keep, a tone that would not have gone out under your name, a field that did not match the system of record.

Keep a short log. Date, what was wrong, whether a person caught it. At the end of the window you can say “the draft missed the access note on two of eight jobs, and the reviewer caught both.” That sentence is more trustworthy than “accuracy improved.”

Do not turn the log into a rate you would not defend to the person who does the job. If the sample is eight instances, you have eight stories and a direction. You do not have a statistic worth putting in a headline.

Operator feedback

Ask the people in the checkpoint one question: would you rather have this draft next month, or not? Then ask what they still fix every time.

Trust is a measure. If the experienced person will not send the draft without rewriting it, you do not have an adoption. You have a demo. Their reasons belong in the decision note. “It sounds like a brochure.” “It forgets the gate code.” “I spend as long checking it as writing it.” Those are design inputs.

A survey with a five-point scale is optional and usually worse than three conversations. You are a small shop. You can talk.

What not to measure yet

Skip revenue attribution, “hours given back to the business,” and any number that requires you to imagine a fully rolled-out future. Those figures are how pilots get oversold and then abandoned when the week does not match the slide.

Also skip vanity activity: messages sent to the bot, seats assigned, logins. Activity is not operation. The adoption strategy cares whether the kept workflow survives contact with the team, not whether people opened the app.

When the numbers are allowed to grow up

After a workflow is in keep, and after more than one person has used it for a month, you can watch the same three signals again: time on the task, caught errors, and whether people still want the draft. That is a trend you earned.

If the kept work now spans tools, needs shared memory, or should show up on a screen the whole shop uses, map tools versus agents versus the human owner on the stack map. Plans for that kind of system are on arcasya.ai. Bring your log to a Free AI evaluation if you want help reading it. Do not bring a target percentage for us to agree with.

Keep going