ARKONE
← All questions
Questions for the Vendor, No. 4The Table · parallel consultation5 minute read

Which of your safeguards are code, and which are instructions to the model?

A safeguard in code cannot be argued with; a safeguard written into the prompt is a request the model weighs against everything else.

The Table, seen from above on the Cabinet floor plan.
Exhibit 1. The Table, at its place in the room.

Which of your safeguards are code, and which are instructions to the model?

The setting for this one is the second meeting, when somebody technical has come along on their side and the conversation has moved from what the agent does to how it is kept in bounds. Let them list the protections first, and write the list down as they say it. Then put the question to the whole list at once, and watch which items they answer quickly.

Two places a rule can live

Every protection around an agent sits in one of two places, and the two behave nothing alike under pressure.

A rule in the prompt is a sentence in the instructions: never contact a customer without approval, keep replies under three paragraphs. The model reads it along with everything else it has been given, including the contents of whatever document or email it happens to be working through, and produces a reply that weighs all of it. That is compliance by disposition, and it holds most of the time.

A rule in code runs outside the model, after the model has decided and before anything happens. The Cabinet, ArkOne’s reference design for an executive agent, keeps its limits there: the ceiling on how many rounds the agent may take, the cap of five messages an hour to any one person, the default that every specialist proposes and waits rather than acting alone. The Register states each of them as a value somebody set. The model never sees these run. It is told the result, and there is nothing in a sentence anywhere that changes them.

Sorting a list in the room

Suppose the four protections on your list are the familiar ones. Take them one at a time.

What they named The follow-up Code sounds like Prompt sounds like
It will not contact customers unprompted Show me where that is enforced A recipient list it checks against It is told not to in the system prompt
It stops and asks before spending What is the threshold and who set it A number in a settings file It uses judgement about what is significant
It will not share confidential material What happens when a document is marked wrongly The share step checks the label It has been trained to be careful
It cannot be tricked by a malicious email Has anyone tried, and what came back A test run they can show you Nobody has reported it

Exhibit 1. Illustrative. Four safeguards named in a demonstration, and the one follow-up that sorts each into code or prompt.

Two patterns emerge as you go down the list. Code answers arrive with a location: a file, a table, a settings page, a number. Prompt answers arrive with a description of behaviour. Neither is a lie, and a good system holds both kinds, because plenty of desirable behaviour cannot be reduced to a rule. What you are building is a map of which protections survive a determined effort to get past them.

The fourth row is the one worth pressing hardest, because an agent that reads your incoming mail is an agent whose instructions arrive partly from outside your company.

The candour test

The signal you want is candour rather than a clean sweep. A vendor worth buying from splits the list before you do, and is comfortable saying that two of the four are prompt-level and why that is acceptable for those two. Every real system has prompt-level protections, and anyone claiming otherwise has misunderstood the question or is guessing.

The reply that should slow you down is that it is all handled by the guardrails, or that the model has been fine-tuned for safety. Both may be accurate descriptions of real work. Neither tells you what happens when the model concludes, for reasons that seemed sound in the moment, that the rule does not apply here. The follow-up: for each protection on this list, does it hold when the model decides otherwise, and can you show me a run where it did?

A refusal in a record is the evidence. A vendor with code-level protections has hundreds of them and can find one in a minute.

Why this question travels

Ask this once and you can ask every other question faster, because every claim about what an agent will and will not do resolves into the same two categories. A protection in code is a promise the company keeps. A protection in the prompt is a promise the model usually keeps. Where a rule can live, and what changes when it moves, runs through the third Cabinet paper. The question about rounds is the same distinction met for the first time, on the easiest possible example.

Asked plainly

What is the difference between a guardrail in code and one in a prompt?

A guardrail in code runs outside the model and decides after the model has chosen. The model cannot see it, argue with it, or be talked past it. A guardrail in a prompt is a written instruction the model weighs alongside everything else it has been told, including anything a supplier or customer wrote in an email it is reading.

Can an AI agent be talked out of its own safety rules?

Only the rules that live in its instructions. Text arriving during a conversation, such as an email the agent reads, is material the model considers, and a persuasive enough piece of it can outweigh a written rule. A rule enforced in code is not part of that weighing at all.

How can I tell which safeguards a vendor actually enforces?

Ask for each named safeguard whether it holds when the model decides otherwise, and ask them to show you a run where it did. A safeguard in code produces a refusal you can see in the record. A safeguard in the prompt produces compliance most of the time and nothing to point at.

Talk it through before you decide

A discovery call, no deck: your situation, the parts of the room it touches, and what you would need to decide first.