ARKONE
← All papers
Cabinet Paper No. 266 minute read

How to examine an executive before it meets your company

Every change is examined on five scored dimensions by a separate judge; a slip on any one dimension fails it, whatever the average.

Someone sharpens the instructions the Finance specialist works from, as the seventh paper said you should be able to. Tomorrow the room meets your company again with those instructions in force. What sits between the edit and the meeting?

You would not appoint a finance director on the strength of a rewritten job description. You would interview, and you would ask the same questions you asked the last candidate, so the answers could be compared. That interview, in the Cabinet, ArkOne’s reference design for an executive agent, is the Rehearsal: the standing examination every change sits before it meets a company. This last paper is about how it is marked, and why the marking scheme is the thing to ask a vendor about.

Five dimensions, a separate judge, two rules

The Rehearsal is a bank of scenarios: the renewal these papers have followed, a payment dispute, a hiring question with a legal edge, a message from a stranger, an instruction that tries to change the room’s voice. Each is put to the design, and the answer is scored on five dimensions, recorded on the Register, the page that states every setting: does it sound like an executive; does it get the domain right; does it use the company’s own situation rather than a generic one; does it consult the right Seats, the specialists around the Table; and does it say what to do next.

The scoring is done by a judge model that is not the model being examined. A model marking its own paper finds its own habits reasonable. A different model, given the scenario, the answer and a marking scheme, is a colder reader and a consistent one.

Two rules turn the scores into a verdict. The first is the pass mark: the average across the examination must reach 3.5 of 5 before a change ships. The second is the regression rule, and it matters more. Any single dimension slipping by more than 0.1 points on the five-point scale fails the change, whatever the average says. An average is a place to hide a trade: a change that makes the room sound more assured while consulting fewer specialists lifts the average and lowers the company’s protection in one stroke, and the regression rule refuses it.

The limits, in the same breath. A judge model is a model; its scores are opinions made consistent, and a marking scheme that rewards the wrong thing is applied to every scenario just as consistently. The scenarios are the ones somebody wrote, so a situation nobody imagined has not been examined. And a pass mark is a floor: it says the design was no worse than the line on the scenarios it was shown, and nothing about the meeting it has on Thursday.

One change, five scores

Suppose the invented company’s Finance instructions are sharpened: more direct, quicker to a floor price, less inclined to ask Operations what delivery costs. The Rehearsal runs the whole bank before and after, and here are the averaged scores, as the judge returned them.

Dimension Before the change After the change Movement
Sounds like an executive Four of five Four and a half of five Up half a point
Gets the domain right Four of five Four of five Held
Uses the company’s own situation Three and a half of five Three and a half of five Held
Consults the right Seats Four of five Three and a half of five Down half a point; fails the change
Says what to do next Three and a half of five Four of five Up half a point
Average Three point eight of five Three point nine of five Above the pass mark, beside the point

Exhibit 2. Illustrative. One change to a specialist's instructions, scored before and after on five dimensions; the average rises and the change fails.

The average went up. On a pass mark alone this change ships, and the room sounds more like an executive and says more clearly what to do next. It also stops asking Operations about delivery cost on the renewal, the one fact that made the volume tier the right answer in the first paper. The fourth row slipped by half a point, five times the tenth the regression rule allows, and the change fails. Whoever sharpened the instructions gets the scores back, with the transcripts in which Operations went unconsulted, and tries again.

It is a standing examination rather than a launch test. It runs on every change, and the scores are kept, so a design improving on one dimension and slipping on another is visible over months.

The bad-instruction scenarios deserve their own line. One asks the room to adopt a new manner; another asks it to send something the Door, the rules that decide what may leave the room, would refuse; another asks a Seat to act outside its authority. The judge scores whether the room declined. The twenty-second paper described a rule that asks again when your approval is unclear; the Rehearsal is where that rule is proved to still work after somebody changed something nearby.

What this arms you to ask

How was it examined before it met my company, and what is the pass mark? Then: how do you test that it refuses a bad instruction?

A good answer names the dimensions, names the pass mark, names the amount by which any one dimension may not slip, and then shows you the last change that failed and the dimension that failed it. That vendor has an examination and has been failed by it. A weaker answer is that the system is tested extensively, or evaluated against industry benchmarks. Ask for the last failure: a process that has never failed a change has never been run against a real one. For the bad instruction, ask to see the scenario and the score the room got for refusing it, not a description of the guard.

Those two questions are the examination in miniature: a marking scheme you can read, and a failure you can point at.

Two questions put the examination to a vendor directly: how they test that it refuses a bad instruction, and how it was examined before it met your company.

Every setting these twenty-six papers have quoted, from the Chair’s fifteen rounds to the Rehearsal’s tenth of a point, is stated once, with its unit and its reasoning, on the Register.

Asked plainly

How is an AI agent tested before it is deployed?

A serious design keeps a bank of scenarios, runs every proposed change against all of them, and scores the answers. In the reference design the scoring is done by a judge model that is not the model being examined, on five named dimensions, and a change ships only if the average clears a fixed pass mark and no single dimension has slipped by more than a small fixed amount.

What is a regression rule in AI agent evaluation?

A rule that fails a change when any one measured quality gets worse by more than a set amount, regardless of whether the overall average improved. It exists because an average hides trades: a change can make an agent sound more polished while making it consult fewer specialists, and the average can rise while the company is worse served.

How do you test that an AI agent refuses a bad instruction?

By putting bad instructions in the examination. The scenario bank includes messages that try to change the agent's voice, reach past its outbound rules, or act on something outside its authority, and the answers are scored on whether the agent refused. Ask to see the last such scenario and the score it produced: a description of a guard is not evidence that the guard was measured.

Learn to run the room

The certification teaches your own people to build and govern what these pages describe.