The deck arrived on Thursday and you opened it once, on a phone, between two other things. Fourteen slides. A photograph of a control room on the cover, then a diagram with your company’s name in a box on the left and the word “agent” in a circle on the right, and an arrow between them going both ways. Somewhere near the back there is a slide headed Security, and it has four ticks on it.
The call is at nine tomorrow. Your head of operations will be there, the finance director, and a person from the vendor whose job title contains the word Solutions. They will show you a screen where somebody types a question in plain English and a tidy answer comes back, and the room will nod, and you will be the one who decides whether that nod meant anything.
Why a demonstration cannot tell you what you need to know
A demonstration is a performance of a chosen case. Someone picked the question, ran it before the call, and watched it work. What you are being shown is the best run, and what you would buy is every run.
The gap between those two is where the decisions live. An agent that answers well, and an agent that answers well and then, unwatched at four in the morning, emails a customer four times about one invoice, look identical for the thirty seconds you are watching. So do an agent that stops when it cannot resolve something and an agent that keeps going until somebody notices the bill.
What separates them sits outside the model and never shows in an answer: a number of rounds after which the meeting must end, a cap on how many messages reach one person in an hour, a list of the people it may address at all, a step where the work stops until a human replies. In the Cabinet, ArkOne’s reference design for an executive agent, those limits have names and values, and each value is written down on the Register, the standing page of every setting in the design.
The questions your team has already prepared will not find any of this. Most of them test capability: can it read our contracts, can it handle Arabic, can it reach our finance system. Every one has a yes available, and the yes is generally earned, because capability is what the last six months of building were spent on.
Limit questions behave differently. They have a number as their answer, or they have nothing. Ask how many times it goes around before it must answer you, and the seller either names a number and where it is set, or does not. There is no impressive way to not know. That asymmetry is why tonight is better spent on twelve questions than on the deck.
Twelve questions, and what each answer is worth
Suppose you take twelve questions into the room tomorrow and ask no others. They are drawn from a longer set, and each is chosen because somebody who has built the thing can answer it in a sentence and somebody who has assembled it cannot.
| The question | A good answer sounds like | What you will often hear |
|---|---|---|
| How many times can it go around before it must answer me? | A number, where it is set, and a run that reached it | It knows when it has enough |
| Which of your safeguards run in code, and which are instructions to the model? | A list of each, kept apart | Both, really, it is layered |
| What sits between it deciding to message someone and the message leaving? | Named checks, in the order they run | There are guardrails around sending |
| How many messages to one person in an hour, and does it know when they are off? | Two numbers and a roster | It is designed to be respectful of that |
| Who may it speak to, and what happens when a stranger emails it? | A roster, and the stranger dropped unanswered | It handles unknown senders appropriately |
| For each part of my business, can it act, propose, or must it ask? | One setting per area, shown on screen | That is fully configurable |
| Which actions wait for a human, and what happens after I answer? | The named actions, and what the run does next | It picks up where it left off |
| Show me the record of the last thing it did and why | A ledger opened live, with the reason beside the action | There is full logging available |
| What does it know about my company tomorrow morning that it knew tonight? | What is written down, when, and where it is stored | It learns as it goes |
| Which model answers which kind of question, and who decided? | Per call, with the reason for each | We use the best model for the task |
| Show me one real invoice for a week of this running, and which line on it you can make smaller | The invoice, and the line, named | Cost depends on your usage |
| How was it examined before it met my company, and what is the pass mark? | Scored scenarios, a mark, and a rule for a slip | It has been extensively tested |
Exhibit 1. Illustrative. Twelve questions, the good answer, and the answer a real vendor gives.
Every good answer contains a noun you could write down: a number, a list, a screen, a file. The right-hand column is made almost entirely of qualities, and a quality cannot be checked in the room or six months later.
None of the right-hand answers is a lie, and this matters more than it sounds. Take each at face value: a limit that is configurable probably is, and a system said to have full logging probably writes a log file somewhere. The trouble is that a true statement about a category tells you nothing about the value. Configurable to what, today, in the instance you would be running. Logged where, and does the system stop working if writing to that log fails.
The three questions in the middle of that table, about who may act and who must ask, are the ones to spend real time on, and they are the reason your finance director should be in the room. Each is a question about authority rather than about software. Somebody in your company will be the person who decides that the agent may issue a credit note up to a certain value on its own and must ask above it. If the vendor cannot show you where that decision is written down and how it is changed, the honest reading is that it has not been made, and it will be made later by whoever is nearest when the question first arises.
Suppose you write four of those answers down as the seller gives them, in a column beside what the reference design would say. The invented company below is a distributor of medical supplies with roughly three hundred staff, and the setting values in the middle column are the design’s own.
| The setting | In the reference design | What to write beside it |
|---|---|---|
| Rounds before it must answer | Fifteen, then the meeting ends | The number, and where it is set |
| Messages to one person in an hour | Five, and none while they are away | Their number, or a blank |
| A near-identical message repeated | Refused inside six hours, reason returned | Refused, or sent again |
| What each specialist may do | Propose and wait, until a person writes otherwise | Act, propose, or escalate, per area |
| When the guard itself breaks | The message goes, the fault is written down | Goes or is held, and who chose |
| The record of an action | Never editable; a correction is a new entry | Editable or not, and by whom |
Exhibit 2.
A blank in the right-hand column is the finding, not a failure to ask properly. Six blanks means the answer to almost every question tomorrow was a quality rather than a value, and you have learned that in an hour rather than in the fourth month of an implementation.
The last row ages best of the six. A ledger you can edit is a set of notes; a ledger you cannot is evidence, and that only becomes interesting on the day somebody disputes what happened. Its limit belongs in the same sentence: a wrong entry stays visible, corrected by a later one rather than removed, and your people should hear that before they see it.
What this arms you to ask
Take three of the twelve if you take nothing else, and take them in this order.
Begin with the safeguards question, because it sorts the room fastest. Ask which safeguards run in code and which are written into the instructions given to the model. A good answer separates the two lists without being pushed. The thinner answer is that the system uses both, and that the prompts are hardened. Your follow-up is one sentence: if I asked the model very insistently to send that message anyway, which of those two lists would stop it. Watch whether the reply names a mechanism or describes an intention.
Then the message question, because it is the one your customers will feel. How many to one person in an hour, and does it know when they are away. In the reference design that is five in an hour, nothing while a person is off or outside their hours, and a near-identical message inside six hours refused, with the reason handed back so the agent reschedules rather than retries; the sixteenth paper walks the whole sequence. Ask the seller to make it try the same send twice inside an hour while you watch, and see whether the second leaves.
Close on authority. For each part of my business, can it act on its own, propose and wait, or must it escalate, and show me the screen where that is set. The reference design ships every specialist at propose and wait; a person writes the permission to act, and that change is itself on the record, which the nineteenth paper sets out. Listen past the agreement and watch for the screen. Your follow-up if none appears: what is it set to right now, in the instance you would be selling me, and who in my company would be able to change it.
Three questions, three follow-ups. If you get through them and the answers have nouns in them, tomorrow was a good use of an hour, whatever you decide afterwards.
That follow-up call is not a refusal, and you should take it. But the person you meet on it is the person you are actually buying from, and the questions in the table are the ones to bring, unchanged, and asked in the same order, so that you can hear where the two sets of answers differ. If they hold up and a contract follows, the terms that decide how it goes are not the ones on the commercial page. You have twelve questions and they fit on one side of a page.
component: null
Asked plainly
What should I ask an AI vendor during a demonstration?
Ask for the limits rather than the abilities. How many times can it go around before it must answer, how many messages will it send one person in an hour, what happens when two copies run at once, and which of its safeguards run in code rather than being written into the instructions given to the model. A demonstration is chosen to succeed; the limits are what govern every run you never watch.
How can I tell a real AI agent from a chatbot with a logo?
A chatbot answers when spoken to. An agent decides to act, uses tools that change things outside itself, and keeps working when nobody is watching. The practical test is whether the seller can name what stops it: a round limit, a rate limit on messages, a list of people it may address, a step that waits for a human, and a written record of every action and its stated reason. If nothing can be named, the thing in front of you answers questions.
Is it a bad sign if a vendor says a setting is configurable?
It is neither good nor bad by itself, and it is the most common reply you will hear. It becomes an answer when the seller says what the value is today, where it is set, who may change it, and whether the change appears in a log. A setting nobody can name the current value of is not being configured by anyone.
