Oversight

The four questions nobody asks in an AI demo⁠.

The demo always goes fine, because the demo is the easy half. Here are the four questions that decide whether an agent is an asset or an exposure, and why most vendors would rather you asked them later.

An empty boardroom, a long table, a row of chairs and a screen at the far end.
Photograph: Office Work by Benjamin Child, CC0 1.0.
01 / The piece

The four questions

We have sat through a lot of AI demonstrations in the last two years, on both sides of the table. They are almost always good. An agent reads an inbox, works out what the message is about, drafts a reply that sounds like a person wrote it, and files the thing in the right place. Nobody in the room is disappointed.

That is not because the vendor is cheating. It is because that part of the problem is genuinely solved, and has been for a while. The model is not the hard part any more. The hard part is everything that happens on the days when nobody is watching, and demonstrations are, by definition, days when everybody is watching.

So we have four questions. They take about ninety seconds to ask and they will tell you more about a supplier than the rest of the meeting put together.

One. Who signed this off?

Not who approved the project. Who approved the individual thing the agent just did. When it sent that email, changed that record, raised that credit note, whose name is against it, and was that person asked before or told after?

There are only three honest answers: a named person approved it in advance, a named person set a standing rule that allowed it, or nobody did. All three can be correct in the right context. What is not correct is a supplier who has never thought about the question, because it means the agent has been built as a tool and not as a member of staff, and tools do not have accountability attached to them.

Our default is that anything leaving your business, spending your money, or changing who can see what, waits for a named person. You have to ask us to loosen that, and we will write down that you asked.

Two. What did it do at three in the morning?

Ask to see the log. Not a dashboard of how many tasks were completed, an actual line by line record of what the agent did, what data it read to decide, which system it touched and what came back.

This matters for the obvious reason, which is that one day something will go wrong and you will need to reconstruct it. It matters more for a less obvious reason: an agent you cannot audit cannot be improved. If you cannot see why it made a decision, every fix is guesswork, and you end up rewriting instructions and hoping.

If the answer is that logging is on the roadmap, the answer is that it is not built. Logging added later is a different product from logging designed in, because the second one records the reasoning and the first one only records the outcome.

The difference looks like this. Note that every line says why, not just what:

The ledger, one decision Reasoning, not just outcome
  1. 02:41:16 Butler Read one message from the shared inbox Read as a delivery query rather than a complaint. Both readings recorded, not only the one it went with.
  2. 02:41:17 Butler Chose the delay wording over the damaged goods wording Because the message names a date and mentions no damage. This is the line that tells you it was not luck.
  3. 02:41:19 Butler Drafted a reply, 14 words changed from the template The 14 words are kept, so a person can see what it wrote and what it merely copied.
  4. 02:41:20 Hardy Held it, because the reply leaves the business Waiting on a named person. No standing exception was configured for this customer.

An example of the format, not a customer’s data. Four seconds of work, and you can reconstruct every step of it eleven months later.

Three. What does it cost to run in March?

Not what the build costs. What the running costs, in a specific month, at your actual volumes, when the busy period lands and the agent does four times as much work as it did in the demo.

This is where a lot of these projects come apart, and it is not usually anybody's fault. The pricing is per unit of work, the units are invisible to the buyer, and the volume is unknown until it is live. Then the invoice arrives and the conversation changes.

The fix is not clever, it is just unpopular: a published monthly price, agreed before you commit, that goes in the budget as a line rather than as an estimate. If a supplier cannot give you one, ask them why the risk of their own pricing model is being carried by you.

Four. What happens the day it sends something wrong to your best customer, and who finds out first?

Every agent will eventually do something you did not intend. That is not a reason to avoid them, it is a reason to plan for it, in the same way you plan for a member of staff having a bad day.

The question is really two questions. Can you stop it, immediately, without ringing anybody? And does the error reach you before it reaches the customer, or after?

A brake that only the supplier can pull is not a brake, it is a phone number. Every agent we build has a stop that you control, and the error path is designed so that a person in your business sees the problem first wherever that is possible.

Why this is not pedantry

Over 40 per cent of agentic AI projects are forecast to be cancelled by the end of 2027, on escalating costs, unclear value or inadequate risk controls.

Gartner, June 2025.

Read that list again, because it is the same four questions in different clothes. Escalating costs is question three. Inadequate risk controls is questions one, two and four. Unclear value is what happens when nobody agreed at the start which number had to move.

None of those are model problems. They are management problems, and they are entirely answerable before a line of code is written. The reason they usually are not is that the demo is more fun than the paperwork, and by the time the paperwork matters the project is already three months old.

Ask them early

The best time to ask these four questions is in the first meeting, of every supplier, including us. A good answer sounds specific and slightly boring. A bad answer sounds like a philosophy.

If you want to see ours in practice, that is what the audit is for: two weeks, fixed fee, and you end up with the four answers written down for your own work rather than for a demonstration.

Back to The Wire

02 / Next step

Tell us the job that never gets done on time.

We will tell you, in plain English, whether an agent can take it, what it would take to build, and what it would cost to run.

Answered by a real person. Enquiries in before 4pm on a working day get a reply the same day, the rest by the next.