Blog

Notes on useful AI, CX, and innovation.

Short working notes on service systems, AI learning, foresight, customer experience, agent workflows, and practical delivery.

Value evidence is the missing step in most AI pilots

AI pilots need a value evidence plan before build starts, so teams can decide whether to proceed, pause, redesign, or scale.

AI pilots often begin with a capability question.

Can the model do the task? Can it answer the question? Can it summarise the case? Can it draft the response? Can it automate part of the workflow?

Those questions matter, but they are not enough.

A pilot can prove that something is technically possible and still fail to show whether the work is worth doing. That is where many AI projects become difficult to judge. They create activity, demos, screenshots, and enthusiasm, but not enough evidence to decide what should happen next.

The missing step is value evidence.

Before a team builds an AI pilot, it should define what evidence would show that the workflow is making the work better.

Capability is not the same as value

AI capability is usually easier to demonstrate than business value.

A model can produce a reasonable draft. It can classify a request. It can search a document set. It can summarise a conversation. It can recommend a next step.

But useful delivery asks a harder question.

Did that support improve the work in a way that matters?

The answer might involve time saved, quality improved, rework reduced, risk made more visible, customer experience improved, employee effort reduced, capacity increased, or decisions made faster.

The point is not to make every pilot heavy or bureaucratic. The point is to make the pilot honest.

If the team cannot say what better looks like, it will struggle to know whether the AI helped.

What value evidence does

Value evidence gives a pilot a decision path.

It helps a sponsor decide whether to proceed, pause, redesign, or scale. It also helps the team avoid treating a working demo as proof of a working business case.

Good value evidence answers four practical questions:

  • What work are we trying to improve?
  • What baseline do we have today?
  • What change would be worth caring about?
  • What evidence will we review before deciding the next step?

These questions do not need a complicated measurement framework at the start. They need enough clarity to stop the project drifting.

For example, if the pilot is meant to reduce manual review time, the team needs to know current review time and what level of reduction would matter. If the pilot is meant to improve consistency, the team needs a way to compare decisions before and after. If the pilot is meant to improve the customer experience, the team needs to know which customer moment is being improved and how that improvement will be noticed.

Without that evidence plan, the team may still learn something, but it will be harder to make a credible decision.

Evidence should fit the work

Not every AI workflow needs the same evidence.

A low-risk internal assistant might be judged on time saved, usefulness ratings, and the quality of human edits. A service workflow might need evidence about accuracy, escalation, customer impact, and operational fit. A decision-support tool might need stronger review of false positives, false negatives, policy alignment, and human oversight.

The evidence should fit the risk, frequency, and value of the work.

That means the team should define the pilot around the actual workflow, not around the tool. Who uses it? What context does it need? What happens when it is uncertain? What decision does it support? What would make the output useful enough for the next person in the process?

Once those questions are clear, the evidence becomes easier to design.

A simple pilot discipline

Before starting an AI pilot, write five lines:

  • The work we are improving is...
  • The current problem is...
  • The value we expect is...
  • The evidence we will review is...
  • The decision we will make after the pilot is...

Those five lines are not a full business case. They are a practical discipline.

They make the pilot easier to sponsor, easier to test, and easier to stop if it is not working.

AI pilots should not only ask whether the system can do something.

They should ask what evidence would prove that doing it matters.

Want to turn this kind of thinking into a workshop, foresight sprint, agent brief, or workflow pilot?

Start a conversation