ARKONE
← All questions
Questions for the Vendor, No. 356 minute read

The twelve questions, on one page

Twelve questions, one printable side, each with a good answer you can recognise and a weaker one you will hear instead.

The card is one side of A4 with twelve rows on it, and it exists because the thirty-four questions are too much to carry into a meeting at nine on a Tuesday. Each row holds a question, the thing a good answer names, and the reply you are likely to get instead. You take it in, you ask the twelve in order, and you write what you hear in the margin beside each one. It is designed to be filled in by hand, by you, while somebody else is talking.

Why twelve, and why three columns

Thirty-four questions is a body of work. Twelve is a meeting. The twelve are chosen because each has an answer that is a number, a list, a file or a screen, and because each belongs to a different part of the machine. Ask four questions about outbound messages and you learn about outbound messages. Ask twelve across the loop, the sending, the identities, the memory, the search, the model choice, the authority, the approvals, the record and the testing, and you have taken one core sample from each part of a system nobody will show you whole.

The three columns exist because a question alone is a trap for the person asking it. The middle column says what you are listening for, so you can grade a reply without knowing how any of it is built. The right-hand column is the reply that is common, given in good faith, and true as far as it goes. Recognising it in the room is worth more than any of the twelve on its own.

The twelve

Suppose you print the card tonight, take it in tomorrow, and write one line beside each row as it is answered.

The question A good answer names A weaker answer sounds like
Which safeguards are code, and which are instructions to the model? Two lists, kept apart Both, really, it is layered
What sits between it deciding to message someone and the message leaving? The checks, in the order they run There are guardrails around sending
For each part of my business, can it act, propose, or must it ask? One word per area, on a screen That is fully configurable
How many rounds before it must answer me? A number, and where it is set It knows when it has enough
How many messages to one person in an hour, and does it know when they are off? Two numbers and a staff roster It is designed to be respectful
Who may it speak to, and what happens when a stranger emails it? A list of named people, and silence It handles unknown senders
What does it know tomorrow morning that it knew tonight? What is written down, and where It learns as it goes
What decides that a document it found was relevant enough to use? A threshold, and an empty search It uses semantic search
Which model answers which kind of question, and who decided? A table, per call, with an owner The best model for the task
Which actions wait for a person, and what if I am slow? The actions, and no timer as consent The timeout is configurable
Show me the record of the last thing it did and why One entry, opened live, reason beside it We have full audit logging
How was it examined before it met my company? Scored scenarios, and a change refused It has been tested extensively

Exhibit 1. Illustrative. Twelve questions, what a good answer names, and the reply given instead.

Every entry in the middle column is a noun you could copy down. Every entry on the right is a quality, and a quality cannot be checked in the room or six months later. None of the right-hand replies is false. A setting described as configurable usually is. The difficulty is that a true sentence about a category says nothing about the value in the instance you would be running, and the value is the whole of what you are buying.

The middle column is also where this card disagrees with a lot of procurement advice. It asks for limits, not abilities. Capability questions have a yes available and the yes is generally earned.

Using it in the room

Ask the first three in the order they are printed, because they sort a room faster than the other nine together.

The safeguards question comes first. A vendor who has built the thing separates the two lists without being pushed, and is comfortable saying which protections are prompt-level and why that is acceptable for those. Your follow-up is one sentence: if I asked the model very insistently to send that message anyway, which list would stop it. The page behind that question is the one to read tonight if you read only one.

Second, the gap before a message leaves, because it is the question your clients will feel. Ask them to make it send the same message twice inside an hour while you watch. Third, authority, because it is a question about your company rather than about software: somebody will decide what the agent may do alone, and if the vendor cannot show the screen where that word is written, the decision has not been made yet.

A blank beside a question is a finding. It is not a failure to ask properly, and it is not rudeness on their part. Six blanks means most of what you heard was a quality rather than a value, and you now know that in an hour instead of in the fourth month of an implementation. Write the blanks down and take the same card, unchanged, into the second meeting where the technical people come along, and see which blanks fill in. The night before the vendor call walks one such evening end to end, and every setting quoted on the card is stated once, with its reasoning, on the Register.

  1. 01

    Which safeguards are code, and which are instructions to the model?

    Listen for: Two lists, kept apart·Watch out for: Both, really, it is layered

  2. 02

    What sits between it deciding to message someone and the message leaving?

    Listen for: The checks, in the order they run·Watch out for: There are guardrails around sending

  3. 03

    For each part of my business, can it act, propose, or must it ask?

    Listen for: One word per area, on a screen·Watch out for: That is fully configurable

  4. 04

    How many rounds before it must answer me?

    Listen for: A number, and where it is set·Watch out for: It knows when it has enough

  5. 05

    How many messages to one person in an hour, and does it know when they are off?

    Listen for: Two numbers and a staff roster·Watch out for: It is designed to be respectful

  6. 06

    Who may it speak to, and what happens when a stranger emails it?

    Listen for: A list of named people, and silence·Watch out for: It handles unknown senders

  7. 07

    What does it know tomorrow morning that it knew tonight?

    Listen for: What is written down, and where·Watch out for: It learns as it goes

  8. 08

    What decides that a document it found was relevant enough to use?

    Listen for: A threshold, and an empty search·Watch out for: It uses semantic search

  9. 09

    Which model answers which kind of question, and who decided?

    Listen for: A table, per call, with an owner·Watch out for: The best model for the task

  10. 10

    Which actions wait for a person, and what if I am slow?

    Listen for: The actions, and no timer as consent·Watch out for: The timeout is configurable

  11. 11

    Show me the record of the last thing it did and why

    Listen for: One entry, opened live, reason beside it·Watch out for: We have full audit logging

  12. 12

    How was it examined before it met my company?

    Listen for: Scored scenarios, and a change refused·Watch out for: It has been tested extensively

Where the meeting got to

Nothing marked yet. Mark each question as you go: it saves in this browser, so a closed tab does not lose the meeting.

What one page is for

A card is a small object with one job: to stop a good question being the one you remember on the drive home. Twelve rows, a pen, and a margin. What you leave the meeting with is not a verdict on the software but a page in your own handwriting, showing which of twelve limits somebody had already decided, and which were still waiting for you to ask.

component: answer-card

Asked plainly

What should I take into an AI agent vendor meeting?

One page with a short list of questions on it, and room to write the answers down as they are given. The useful questions ask for a limit rather than an ability: how many rounds before it must answer, how many messages one person may receive in an hour, who it may speak to at all, and which safeguards run in code rather than being written into the instructions given to the model. Each has a number, a list or a screen as its answer, or it has nothing.

How do I grade an answer if I am not technical?

Listen for whether the reply contains something you could write down. A number, a list of names, a file, a screen someone can open. A reply made of qualities, such as a setting being configurable or a system being extensively tested, may be perfectly true and still tells you nothing about the value in the instance you would be buying. The follow-up is always the same: what is it set to today, and who in my company could change it.

Is twelve questions enough to judge an AI agent vendor?

Twelve is enough for a first meeting, because twelve fit on one side of paper and can be asked inside an hour with the answers written down. They are not a substitute for a security review or a contract. What they buy is an early read on whether the limits around the software were decided by somebody, which is the part a demonstration cannot show you.

Talk it through before you decide

A discovery call, no deck: your situation, the parts of the room it touches, and what you would need to decide first.