The chatbot may attract the reviewer’s attention, but the regulated activity begins where the conversation turns into a transaction. If the product helps a user move money, obtain credit, buy an investment or complete another controlled action, the compliance boundary follows that action through every handoff.
In 2017, Uber’s software presented regulators with a problem that could not be understood by looking at the app alone. The New York Times reported that Uber used a system called Greyball to identify some government officials and show them a different version of the service, including in Portland, Oregon. A person could open the same app as everyone else and see something that did not reflect the operation behind it.
At that point, the screen stopped being reliable evidence of what the system did.
The screenshot showed where the product changed roles
The Nigerian founder entered the review prepared to explain the chatbot.
He could describe the model, the prompts and the limits placed on its answers. The bot gathered a customer’s intent, asked follow-up questions and presented possible next steps. That was the part his team had built, tested and demonstrated.
Then the reviewer opened a WhatsApp screenshot.
The customer had asked the bot what to do next. The bot had collected enough information to identify an option, then directed the customer toward a human-assisted process outside the chat. That handoff was the important part. Somewhere after the friendly conversation, a person or connected service helped complete an activity subject to rules the chatbot itself could not avoid.
The review therefore moved away from model behaviour. The questions became operational.
Who received the customer’s details? What could that person do with them? Who made the final decision? Where was consent recorded? Could the customer tell when they had left an automated conversation and entered a regulated process? What happened if the advice in WhatsApp differed from what the downstream operator eventually offered?
A polished answer about the chatbot could not settle those questions.
Regulation follows the completed job
Founders often define a product by the interface they built. Regulators are more likely to examine the job the customer can complete.
That difference matters for AI products because the interface can appear several steps removed from the consequential action. A bot may “only” collect information. An agent may “only” recommend an option. A workflow may “only” send the user to a partner.
Each description can be technically accurate while hiding the full chain.
The practical review starts with the customer’s first message and ends after the real-world outcome. Draw every step, including the ones owned by staff, contractors, partners and tools your team did not build. Mark where data changes hands, where money moves, where eligibility is assessed and where someone can approve, reject or alter the result.
This is the same discipline required when a demo depends on an operation sitting outside the product. In the shared spreadsheet the demo did not account for, the hidden manual process changed the meaning of what customers were being shown. Compliance reviews expose the same gap, with higher consequences.
The off-ramp deserves special attention because that is where a conversational product often becomes something else. A chatbot answering general questions creates one risk profile. A chatbot gathering personal information and routing a customer into a credit, payment, insurance or investment process creates another.
The wording around the handoff does not decide which one you have built. The actual workflow does.
Audit the handoff before polishing the model
A useful internal review can begin with one customer conversation.
Choose a real completed journey, remove personal information and reconstruct what happened after every message. Do not stop when the bot produced its final response. Follow the link, notification, staff action, spreadsheet entry, partner request and customer outcome.
Then ask who had authority at each point.
If a staff member could change the offer, record that. If the partner applied its own checks, document where those began. If the customer could not tell which company they were dealing with, fix the disclosure. If consent lived only in a chat message that nobody could retrieve consistently, treat that as an operational gap.
Permission controls matter here too. An AI agent that can trigger an external action needs a clear boundary around what it may attempt, what requires confirmation and how access can be withdrawn. Mensah’s failed purchase as proof of permission revocation shows why a settings screen is weaker evidence than a blocked action.
The cheapest time to find these gaps is before a reviewer finds them in a customer screenshot.
The interface cannot carry the whole defence
Uber’s Greyball system became significant because the visible interface did not tell officials the whole truth about the operation. The New York Times report focused attention on the gap between what appeared on a phone and what the company’s system was doing behind it.
An AI founder faces a smaller and different decision, but the mechanism is familiar. A reviewer will look past the conversational layer when the customer’s journey continues into an activity with legal consequences.
Before the next compliance meeting, print one real conversation and trace it to the final human or system action. Circle the moment the product changes roles. That circle is where the review should start.
Comments
No comments yet.