ARKONE
← All questions
Questions for the Vendor, No. 26The Door · the outbound guardrails4 minute read

When a rule is broken, what is the model told?

A guard that refuses in silence teaches the agent nothing, so it tries the same thing again until something gives way.

The Door, seen from above on the Cabinet floor plan.
Exhibit 1. The Door, at its place in the room.

When a rule is broken, what is the model told?

Somebody shows you a log from the pilot. One action appears in it a dozen times inside four minutes, each attempt stopped, each identical to the last. Nothing was misconfigured. The agent asked to do something, was refused without being told why, and did what a diligent worker does when a task fails for no visible reason. Your guard held, and it also spent your money in a circle.

A refusal is a message, or it is a wall

The guards deciding what may leave an agent are called the Door in ArkOne’s reference design for an executive agent, the Cabinet: rules in code, run after the model has decided to send. They check the recipient, the hour, the volume and whether the same thing has already been said.

What matters here is the direction the answer travels. In the reference design a refusal is handed back as a sentence naming which rule refused and why, a setting on the Register. Told that a near-identical message went to the same person this morning, the agent can wait, write something genuinely new, or take the matter to a person.

Told nothing, it has no such option. An unexplained failure and a temporary one look the same from inside, and one more attempt is the reasonable response to a temporary one. The wall does not teach. It absorbs, once per attempt, at the price of a turn. The paper on the Door has the six rules.

What a refusal turns into

Suppose you carry the log into three meetings and ask what its dozen identical lines would have looked like on their product.

What a good answer sounds like What you will be offered instead What that hides
It gets a sentence back naming the rule and the window, and it reschedules rather than repeats The action is blocked and the block is recorded A record written for you, not a reply written for the agent. Both can be true, and only one changes what happens next
The refusal says what would make the action acceptable, so the next attempt differs from the last It receives a standard error An error with no reason in it. Ask what the text of that error says, out loud, and count the words
A repeated attempt against the same rule raises an alarm, not a third quiet refusal Our guardrails are enforced at the platform layer Where enforcement happens, which was not the question. A rule can hold perfectly and still hand back nothing

Exhibit 1. Illustrative. Three replies about refusals, and what each of the weaker two turns a no into.

The first row passes most procurements, because a block that is written down satisfies the instinct that something must be auditable. Auditing is retrospective by nature, and the sentence returned to the agent is the part of the refusal that alters the next four minutes.

The third row deserves a slower reading. Suppose a guard refuses correctly every time and says nothing each time. Nothing is ever sent, which is the outcome you wanted, and nothing is reported either, because from the guard’s side each refusal is a success. What you buy in that arrangement is a silent tax, paid per attempt, found by somebody who opened a log for a different reason.

The demonstration that settles it

Ask them to break a rule in front of you and read the reply aloud.

Pick the easy one: have it send the same person the same message twice within the hour. The first goes, the second should stop, and what you want on the screen is the exact text handed back to the agent at that moment, not the entry in the dashboard. Listen for whether it names a rule and a condition, or whether it is a number and a word.

Where a sentence comes back, ask what the agent does next. Waiting is a good answer, and so is taking the matter to a person. Attempting a slight rewording is not, because a guard that can be talked past by rephrasing protects against accidents rather than against intent, and both of you should be plain about which is on offer. That sits alongside whether a rule is code or an instruction.

Where nothing comes back, ask what stops an attempt repeating, and listen for a limit outside the model: a retry ceiling, a lengthening pause, an alarm on the third try.

Every other guard question asks whether the no holds. This one asks what the no says, and the difference surfaces on the bill, not in the incident report. A rule that only stops things is half a rule.

component: answer-card

Asked plainly

What happens when an AI agent is blocked by a safety rule?

That depends entirely on what the blocking code hands back. If it returns a plain sentence naming the rule and the reason, the agent can change course: wait, rephrase, or take the matter to a person. If it returns nothing, or a bare failure code, the agent reads an unexplained failure and its most reasonable next move is to attempt the same thing again.

Why does an AI agent repeat an action it has already been stopped from taking?

Because a failure with no reason attached looks like bad luck rather than a decision. A model handles what looks like a transient fault the way any careful worker would, by trying once more. Repeated attempts against a silent guard are usually a design fault in the guard, not stubbornness in the model.

Should guardrails explain themselves to the AI agent they are stopping?

Yes, and the explanation is written for the agent, not for a dashboard. A refusal that names which rule refused, and what would make the action acceptable, converts a blocked attempt into a sensible next step. It also gives the people reading the audit trail a sentence in plain words rather than a code to look up.

Talk it through before you decide

A discovery call, no deck: your situation, the parts of the room it touches, and what you would need to decide first.