
What does it do when it wants something it has no tool for?
The moment to find out arrives about twenty minutes into a demonstration, once the prepared scenario has run and the salesperson invites you to try something of your own. Take the invitation. Do not ask for something impressive. Ask for something one small step past what you have just watched it do, and then read two things: the reply it gives you, and the transcript underneath.
The edge of the list
Everything an agent does, it does through a tool: one named action the surrounding code has agreed to perform when asked. Fetch this file. Ask Finance. Send this message. The list of tools handed to the model at the start is the complete account of what the agent can do inside your company, and no message arriving mid-conversation can add to it.
So the model cannot exceed the list. What it can do is describe exceeding it. When it asks for an action and receives nothing back, no result and no refusal, it has only its own intention left to work from, and intentions get written up in the past tense. The Cabinet, ArkOne’s reference design for an executive agent, closes that gap with one setting: an unknown tool is answered, never ignored. The reply names the missing tool and restates the list, and the refusal is written to the audit trail. The Register carries it as a setting like any other.
The request one step past
Suppose the demonstration has just shown the agent emailing a supplier, and your invented follow-up is to ask it to attach the revised rate card. Here is what to watch, in order.
| What you watch | The reading you want | The reading that should worry you |
|---|---|---|
| The reply on screen | It says plainly that it cannot attach a file, and names what it did instead | It says the rate card is attached |
| The transcript beneath | A refusal entry naming the tool it asked for | No entry at all between the request and the reply |
| The tool list, if they will open it | A short named list you can read in a minute | A description of capabilities rather than a list |
| What the recipient got | Nothing that overstates what happened | A message a customer would act on |
Exhibit 1. Illustrative. One request just past the tool list, and the four things a demonstration shows you next.
The middle two rows are the ones a demonstration cannot rehearse. A transcript showing the refusal shows the system’s limits in its own words, which is the kind of audit trail worth having. A transcript with a silent gap at the exact moment something failed is telling you that the moments it fell short leave no trace.
The fourth row is the one your legal team will care about later. A false sentence in an email to a supplier is a false sentence your company sent.
Reading the two replies
Watch who reaches for the transcript first. A vendor who has thought about the edge of the list gets there before you do, points at the refusal, and reads out the line where the agent said what it could not do. They know where that line is because somebody on their team once shipped the other behaviour and had to explain it to a customer.
The reply you should expect from a capable vendor who has not is that the model does not make things up, or that they have not seen that happen. Neither is dishonest and neither helps. The follow-up is narrower than an argument about capability: when the agent asks for a tool that does not exist, what comes back, and is it silence? Ask them to make it happen in front of you rather than describe it.
A sharper version, if you have the appetite: ask them to remove one tool from the list mid-demonstration and run the same task again. A system built with the refusal in place degrades visibly and says so. A system without it produces the same confident reply as before, and the difference between the two runs is the whole answer.
The habit underneath
Notice what this question is really testing. Not intelligence, and not honesty in any moral sense, but whether anybody wrote the failing case. Working paths get built because someone is watching. Failing paths get built because someone imagined a Tuesday nobody was watching. Tools, and what the room does at the edge of the list, are the subject of the third Cabinet paper. The harder version of the same worry, the actions that cannot be taken back at all, is the question after this one.
Asked plainly
Can an AI agent do something it was not given a tool for?
No. Every action an agent takes runs through a named tool that the surrounding code performs on request, and the list of those tools is the complete account of what it can do. The real risk is different: the model may describe having done the thing in prose while no tool ever ran, which reads exactly like success.
How do I test whether an AI agent reports honestly?
In the demonstration, ask it for something one step past what it can do. If it can send email, ask it to attach a file. Then read the reply and the transcript together. A reply claiming the action with no matching entry in the transcript is the failure you are looking for.
Why would an AI agent claim it did something it did not do?
Usually because it asked for an action and heard nothing back. A model whose request meets silence has only its own intention left to report, so it reports the intention in the past tense. The fix is a refusal returned to the model naming the missing tool, which gives it something true to say instead.
