Alfred AnyanInsights
← All insights

AI Demo Reliability: What a No-Rescue Test Taught Kojo About Human Review

Two people working together on a creative project, reviewing photos on a laptop indoors.

Photo by Ron Lach on Pexels

A convincing AI demo can hide a founder doing the work the product appears to do alone. The real test begins when every successful result must survive without the founder correcting inputs, rewriting outputs or quietly choosing the safest path.

At 10:47 on Sunday night, Kojo was replying to another message praising the demo. On his desk in Accra, a cold cup of coffee sat beside two browser windows: the polished result his prospect had seen on Friday, and the prompt history showing how many times he had intervened to produce it.

Kojo is an invented composite, but the decision is familiar. His prospect wanted to run a pilot on Monday. If he accepted, a bad result could cost him the contract. If he delayed, the prospect might choose another supplier.

He typed, “Glad it landed well,” then stopped. The demo had landed because Kojo caught every weak answer before anyone else saw it.

The work hidden behind the result

During Friday’s call, the prospect uploaded a clean document. Kojo had already tested that format. When the model misunderstood one field, he rephrased the prompt while explaining another screen. When the output included an uncertain recommendation, he removed the sentence before sharing his window.

None of this required dishonesty. Founders guide demos all the time. The problem was that Kojo had started treating his own judgment as part of the product.

I have made versions of this mistake while building software. When you know the intended result, you compensate without noticing. You choose the file most likely to work. You skip the screen that still breaks. You recognize a wrong answer because you understand the domain, then correct it before the customer can react.

Each intervention makes the demo stronger and the evidence weaker.

That distinction matters when runway is short. A founder can spend another month improving the model because the demo looked promising, only to discover that customers cannot reproduce the result. The alternative can be equally expensive: hiring someone to sell a workflow that still depends on the founder sitting behind it.

The unresolved question is simple: did the product complete the job, or did the founder complete it through the product?

A Sunday test with no rescue

Kojo closed the polished browser window and created a test he could not steer. He took a document he had not prepared for the demo, wrote the instructions a new customer might reasonably write, and promised himself he would not edit anything until the process ended.

The first output looked confident and missed a key constraint.

He ran it again. The second output changed the wording but kept the same error. By then, accepting the Monday pilot felt reckless. The prospect would be making a real decision from an answer Kojo knew could fail quietly.

This is where founders often reach for more prompt work. Sometimes that helps. Sometimes it extends the performance by teaching the founder another way to rescue the result.

A better next step is to record every rescue. Write down where you changed the input, selected between outputs, supplied missing context or discarded an answer. That list describes the operational role the product still expects someone to perform.

It also tells you what to build next. One intervention may need a product control. Another may require a narrower promise. A third may belong in a paid human review step because the model cannot yet handle it reliably.

The same issue appears when a demo avoids a difficult screen. I wrote about that trade-off in the missing page your AI demo avoided. The avoided step often contains the customer’s real risk.

The decision behind the demo

Just after midnight, Kojo changed his reply. He offered a limited pilot with human review stated plainly, a narrow document type and a checkpoint before any output informed a customer decision.

That choice made the product sound smaller. It also made the promise true.

He did not need to pretend the AI worked alone. He needed to decide which part customers could trust now, which part required his hands and whether they would pay for the combined service while he improved the software.

This is especially important across markets. A workflow tested with familiar documents in Accra may meet different language, formatting and operating assumptions in Lagos, Berlin or Atlanta. Geography does not automatically break the model. It exposes which assumptions the founder had been supplying from memory.

Pausing a fintech launch when demo conditions did not match customers explores the same decision at a higher-risk boundary. The useful move is to compare the conditions of success, not the confidence of the presentation.

What to count before Monday

Before showing the next prospect, run the workflow once without rescuing it. Count each moment when you feel tempted to explain, edit, retry or choose a better input.

Then classify those moments. Which intervention can the product remove? Which must become part of the service? Which reveals that the current promise is too broad?

On Monday morning, Kojo opened the pilot call with the review step visible on the screen. The prospect saw where the AI stopped and where Kojo’s judgment began. The demo was less magical than Friday’s.

For the first time, it could survive his absence.

Comments

No comments yet.