An AI launch should be postponed when conflicting customer records make its outputs impossible to verify. Better model performance cannot rescue a product whose source data gives several answers to the same question.
Consider Chidi, an illustrative composite of a Nigerian founder. At 8:17 on a Monday morning in Lagos, he was holding a paper cup of coffee and watching his team’s launch dashboard refresh. Their AI assistant was due to go live that afternoon, after weeks of prompt tests, model comparisons and late-night fixes.
Then the same customer appeared three times.
One record showed an active account. Another marked it overdue. A third used an old company name and listed a payment that the finance export could not find. When Chidi asked the assistant for the customer’s status, it produced a confident answer based on whichever record entered the retrieval process first.
The wording sounded polished. The answer could still be wrong.
The model passed while the product failed
The team had spent Friday clearing the model for launch. It followed instructions, used the requested tone and handled the test questions well. On the evaluation sheet, it looked ready.
Monday exposed the missing test: could the team trust the facts beneath each response?
Chidi tried another account. The assistant called it a renewal opportunity, while the billing system showed an unresolved balance. A third query returned a contact who had left the customer’s company months earlier. Each answer was plausible enough to reach a sales call, an account review or a management report without raising immediate suspicion.
That made the problem dangerous. An obviously broken answer invites correction. A fluent answer built from conflicting records earns trust before anyone checks it.
By 9:05, the bad ending was clear. If the team launched that afternoon, the assistant could give staff incorrect customer information at the exact moment they needed certainty. One wrong account decision could damage a relationship the company had spent years building.
The launch message was already drafted. Internal expectations were set. Chidi postponed it anyway.
AI exposes the decisions hidden inside your data
Teams often treat data cleaning as preparation for the interesting AI work. In practice, AI turns old data disagreements into live product behavior.
A salesperson may know that “M. Okafor Ltd” and “Okafor Manufacturing” refer to the same account. Someone in finance may remember why two totals differ. An operations lead may recognize which spreadsheet contains the latest status. That knowledge can keep a manual process moving because people quietly repair the gaps as they work.
A model has no access to those unwritten corrections. It sees records, fields and retrieval rules. If two systems disagree, the product needs an explicit policy for deciding which source wins, when to ask for review and how to show uncertainty.
This resembles the control problem in what two credible totals taught Ama about revenue reconciliation. Two numbers can each have a reasonable origin while still leaving the business unable to act. The real work begins when someone assigns ownership, defines the source of truth and records how exceptions will be resolved.
AI makes that discipline visible. Every generated answer carries an implied claim: these are the facts the business accepts.
A trustworthy launch starts with record-level tests
Chidi’s team stopped scoring only the final response and traced each answer back to its supporting records. They created a small set of high-risk customer questions and inspected the full path from source system to generated text.
For every answer, they asked:
- Which record supplied this fact?
- Does another system disagree?
- Who owns the decision when records conflict?
- Should the assistant answer, qualify its answer or refuse?
- Can a staff member see the evidence before acting?
The exercise changed the launch criteria. Accuracy now included traceability, conflict handling and a clear route to human review. A graceful refusal became more valuable than a polished guess.
The team also found knowledge that lived in people rather than systems. One account manager could explain several naming mismatches from memory, but none of those explanations had been documented. That is the same fragility explored in Kweku’s company runs on memory: a process can appear functional while one person silently holds the map.
With the launch paused, those hidden decisions became rules the product could apply consistently.
The cleared model was only one component
Later that week, Chidi returned to the same dashboard. The duplicate customer records had been linked, the billing source had clear authority for payment status, and disputed fields triggered review instead of generating certainty.
He typed the question that had stopped the launch.
This time, the assistant showed the current account status, identified the source and flagged the remaining disagreement for a person to resolve. The response was less impressive than Monday’s fluent paragraph. It was far more useful.
That is the durable standard for an AI product: the team can explain where an answer came from, what could make it wrong and who decides when the underlying facts disagree.
Before clearing your next model, choose ten questions that could cost your company money, trust or time if answered incorrectly. Trace every answer to the original record. Where the trail breaks, keep the launch button untouched.
Comments
No comments yet.