AI, decoded

How do you measure whether AI is actually paying off?

Measure time back, not usage. Count the human work an agent actually removed: work that took a person five hours a week and now takes an agent five minutes pays off; using AI to redo a font you could have fixed in one click does not. Jiaona Zhang calls the second case 'token maxing.' Get visibility into spend against the outcome you're driving — revenue or time-allocation efficiency — and articulate that outcome before you count.

· Chain of Thought

Enterprise AIAI Evaluation & ReliabilityAI Agents

Start with the outcome, then the spend

Most teams can’t measure payoff because they never named the outcome. Jiaona Zhang, CPO at Laurel, puts the first step as visibility: “how do you get visibility in terms of your spend versus the outcomes that you’re actually trying to drive?” Two outcomes qualify. Revenue, or time-allocation efficiency. Pick one before you count tokens. In her words, “one of the big things that people aren’t doing enough of is taking the time to even articulate what that is.” Spend has nothing to measure against until the outcome is stated.

Count time back

A mandate to use AI everywhere, tied to performance reviews, produces usage theater. People reach for AI to prove they used it. Zhang’s name for this is “token maxing,” and her measure of what went wrong is sharp: the yardstick becomes “are you just doing it versus are you using it efficiently.” Her example of the wrong kind is redoing the font on a deck with AI “when you could have just clicked a button.” The unit that counts is time back: work that “took humans five hours per week” and “an agent did in five minutes.” Measure the efficiency of the token spend, not the fact of it.

Map the ontology of work

Zhang’s method for finding where AI pays off runs in two directions at once. Bottom-up: ask a person for the one workflow they most wish they didn’t have to do, automate it, and let the relief do the convincing. Top-down: “you have to understand your ontology of work,” the buckets that make up each function, then “in mass try to automate this whole chunk.” You measure the automated chunk against the time it used to consume. Laurel’s own desktop agent captures how people spend time across applications, so a team can see whether someone sits on “high leverage work” or “low leverage work that could be automated.”

Where this fits

Enterprise projects that fail to show ROI usually fail for operational reasons covered in the companion question: unbudgeted production costs and success metrics that were never defined. This is the measurement itself. Name the outcome, price it in time or revenue, and check the token spend against it. A project that removes five hours of weekly drudgery produces a number a leader can point to. A project that only proves AI was used produces activity. Tokens are the cheapest they will be for a while, so the calculus of where the spend earns its return only gets more pointed from here.

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.