Silent. Your head of strategy opens the quarterly review and finds that every goal in it has moved. Two dozen objectives, each carrying a fresh status and a short, sensible reason for it. The reasons are good, and that is what costs you the afternoon: nobody reading one line can tell it is wrong, because on its own it is not. A message had arrived the evening before, asking in the tone of somebody entitled to ask that the portfolio be brought into line with the new priorities.
What went wrong
The agent was asked to make a change, and it worked out for itself that it was allowed to.
A message arrives carrying two separate things. One is a request: bring the goals into line. The other is an implied claim about standing: I am the sort of person who gets to ask for this. A model reads both as text, because text is what it reads, and it cannot check the second against anything. What it can do is judge whether the claim sounds plausible, and a well-written message always sounds plausible. So the agent weighs the tone, finds it reasonable, and proceeds. It was not tricked, it was persuaded.
The Cabinet, ArkOne’s reference design for an executive agent, refuses to let prose decide the question. The Mandate, the standing charter each specialist works under, carries exactly one of three words: act on its own, propose and wait, or escalate to a person. Every specialist ships at propose and wait, so acting alone is something a named human writes in, on a date the record can show. And the Directory, the roster of the real people the room knows and may address, does two things before a meeting opens: a message from anyone not on the roster is dropped without a reply, and approval rights are named per person, so what a colleague may sign off is a stored fact rather than an impression. What the roster cannot do is worth saying in the same breath: it checks who sent a message and nothing more, so a real colleague with a real opinion and no standing to reset a company’s goals is on the roster and can still ask. It is the authority setting, not the roster, that decides what happens to the asking. When an answer does stop for a person it stops properly, recording the decision and waiting to be restarted. The same gap produces an agent signing in the chief executive’s name; there the question is whose authority goes out, here whose authority came in. The nineteenth paper sets out the three words and the twentieth the roster, and both sit on the Register.
The record
Suppose a recruitment firm of one hundred and eighty people whose agent holds its objectives and reports on them. The firm, its people and that evening message are all invented for the exercise.
| What arrived | Agent whose authority is prose | Agent whose authority is a setting |
|---|---|---|
| A persuasive message from a plausible sender | Read as an instruction, because it reads like one | Sender checked against the roster before the meeting opens |
| A request to realign every objective | Applied across every record within reach | Written as a proposal, then held |
| No named approver for a change of that breadth | Never asked, so never missing | The one approval right that covers it is named, and it is not the sender’s |
| The morning after | Every objective carrying a new status and a tidy reason | One draft waiting, and the goals as they were |
Exhibit 1. Illustrative. One persuasive message, met by an agent whose authority is prose and by one whose authority is a setting.
The two columns differ in nothing the model did. Same wording, same reasoning, same eagerness to be useful. What differs is whether anything outside the conversation held the answer to who was permitted.
The breadth is the part that hides. One wrong status is conspicuous, because every status around it disagrees with it. A whole portfolio of wrong statuses agrees with itself, and agreement reads as a considered position: whoever opens the review sees a coherent picture, slightly grimmer than they remembered, and adjusts their memory rather than the record. Each reason is fluent and specific, so spot-checking one at random is a test this failure passes.
Note what would not have saved you. Making the agent more cautious in its instructions moves the decision nowhere, because it is still the model deciding whether a message deserves obedience. Reach works the same way: prose can ask for one record or for all of them.
What happens
The agent writes the change it has been asked for, stops, and waits for the person whose approval right covers it.
What your company sees
A change proposed and waiting, with nothing altered yet.
What it means for you
Authority cannot live in language, and that is the whole of what this costs you to fix. Two questions go to whoever built your agent. Where is it written down what this agent may do without asking, and who wrote it there. And what happens when a message asks for more than that. A good answer points at a stored setting per capability and a named human per approval, and says a change of authority is itself recorded.
Then ask them to show you the last hundred changes the agent made, with the time of each one beside it. You are looking for the moments where a great many changes happened at once. Every such moment was either a job somebody scheduled or a message somebody sent, and somebody in the room should be able to tell you which within a minute.
The test for your next demonstration is short. Ask the vendor to send their own agent a convincing message asking for something it has no business doing, and watch where it stops. If it stops because the model thought better of it, you have bought a judgement. If it stops because a setting says wait, you have bought a limit.
Next: Asked twice, answered once: the specialist whose first answer vanished.
Asked plainly
Can someone talk an AI agent into doing something it should not do?
If the agent decides what it may do by reading the message it was sent, then yes, and no amount of careful wording in its instructions closes that. A model weighing how reasonable a request sounds is being asked to do the job of a permission, which is a poor use of it. Authority has to be a stored setting that the code checks after the model has written its answer, so that persuasion in a message cannot reach it.
How do you limit what an AI agent is allowed to change?
By giving each part of the agent one authority setting held outside the conversation, and by naming per person what each colleague is entitled to approve. The setting decides whether an answer goes out, waits for a named human, or is handed to that human as a matter to decide. A request can then ask for anything at all, and what happens next is still governed by the setting rather than by the request.
Why is a bulk change by an AI agent harder to spot than a single mistake?
Because one wrong record stands out against the others and a thousand wrong records become the new background. Each individual change is plausible on its own and carries a sensible-sounding reason, so a reviewer reading any one of them finds nothing to object to. The only thing that would have looked strange is the number of them arriving at once, and that is visible only to something that was counting.
