The cost of putting an AI agent in front of your customers has collapsed. Capable models are rented by the token. Orchestration frameworks are open source. A vendor can stand up a working demonstration in days and a pilot in weeks. What once required a research team and a year now requires a subscription and a sprint.
The cost of trusting that agent has not moved. Trust is not a property of the software. It is a judgment about how a system will behave with real users, real data, and real stakes – and it is only as good as the evidence behind it. Gathering that evidence still takes time, method, and a reviewer with no interest in the answer coming out a particular way.
That gap – between how cheap deployment has become and how expensive trust remains – is where most agent failures now begin. Organizations close it in one of two ways: with evidence, or with hope. This note is about the first way.
Why the demo always performs
A vendor demonstration is not deceptive, but it is curated. It runs on stable connectivity, with clean inputs, in the language the system handles best, guided by a presenter who knows which questions to ask. Benchmarks have a similar shape: they measure a system against the distribution it was tuned on, and vendors naturally optimize for the tests they know buyers will check.
Real users do not behave like benchmarks. In much of Africa, they arrive over WhatsApp on intermittent connections. They switch between Swahili, English, and Sheng within a single message. They pay through M-Pesa and mobile money rather than card rails, abbreviate heavily, and ask questions no product team in another market anticipated. A system tuned elsewhere meets all of this at once. Performance that transfers cleanly across that distance is the exception, not the default — and nothing in the demo tells you which case you are buying.
The asymmetry of risk
A multinational can absorb a failed agent deployment. It writes off the pilot, retires it quietly, and moves on. A growing African firm faces different arithmetic. Its reputation is concentrated in a smaller customer base and travels fast through tight networks. Its legal exposure is real: an agent that mishandles personal data engages obligations under Kenya’s Data Protection Act, 2019 – obligations that turn on what is done with the data, not on the size of the firm doing it. And the harm lands on customers who often have fewer alternatives and less recourse.
The vendor’s risk is not your risk. A vendor sells the same agent into many markets; a failure in yours is a support ticket to them and a material event to you. This asymmetry is the strongest argument for independent scrutiny before deployment: the party carrying the most risk should not depend for its evidence on the party carrying the least.
The limits of marking your own work
Internal quality assurance matters, and good teams take it seriously. But internal teams cannot fully audit their own systems, for reasons that have nothing to do with competence. They built the thing, so they test the paths they designed and struggle to imagine the ones they did not. They carry the deadline, so findings that threaten it face a headwind no one has to articulate. They see what they expect to see. Every discipline that takes verification seriously – financial audit, aviation safety, clinical research – separates the people who build from the people who check, for exactly these reasons.
Independent review is not internal QA repeated by different staff. It produces things internal testing structurally cannot:
- Adversarial testing that probes for the inputs and failure modes the team did not design for, rather than confirming the paths it did.
- An affected-community perspective – evaluation grounded in the languages, platforms, and conditions of the people who will actually use the system.
- Documented reasoning – findings and evidence that a board, a regulator, or a commercial partner can examine, rather than a verbal assurance that testing happened.
- Public-facing summaries, where appropriate, suitable for disclosure to the communities the system affects.
The measure of a good review is not reassurance. It is findings specific enough to act on.
Regulation is moving toward evidence
The regulatory direction is consistent across jurisdictions. The EU AI Act establishes assessment and documentation expectations for higher-risk systems, and those expectations travel: African firms that export to European markets or sit in European supply chains will encounter them in contracts and due-diligence questionnaires whether or not they are directly regulated. The African Union’s Continental AI Strategy signals a continent-level commitment to AI governance. Kenya’s National AI Strategy points in the same direction at national level, alongside the data protection regime already in force.
None of these instruments needs to be predicted in detail to see what they share: a shift from asserted responsibility to demonstrated responsibility. Organizations that build the habit of independent, documented review now will meet those expectations as routine. Organizations that wait will meet them as a crisis, on someone else’s timetable.
Where to begin
Adili AI conducts independent ethical review of AI systems, including agents built by vendors and agents built in-house. Each review is scoped to the system, its context, and the decision it needs to inform – a procurement choice, a pre-launch check, or a post-incident examination. We keep the two sides of our practice separate: we do not review systems we have built, a boundary described in our responsible AI development work.
Not every agent needs a full review, and we will say so when it does not. A scoping conversation takes an hour and settles the question of proportionality: what could this system get wrong, who would carry the cost, and what evidence would settle the matter. If you are weighing an agent deployment – or already living with one – contact us to arrange that conversation.