ARKONE
← All papers
Cabinet Paper No. 9The Brief · the standing instructions6 minute read

What the Chair reads before every meeting, in the order that decides your bill

An agent re-reads its whole briefing every turn, so what changes least goes at the top, and that order decides the bill.

The Brief, seen from above on the Cabinet floor plan.
Exhibit 1. The Brief, at its place in the room.

Your first month’s invoice for the agent arrives, larger than the demonstration suggested. The price per word looked modest, and nothing in the conversations you read explains the gap. The explanation sits in a part of the system nobody showed you: the briefing the agent reads before every single turn.

That briefing has a name in the reference design this series describes. The Brief is everything handed to the Chair, the one voice that runs the room and speaks for it, before the Chair says anything, and the design it belongs to is the Cabinet, ArkOne’s reference design for an executive agent. You pay for every word of the Brief, on every turn, and what you pay depends less on its length than on the order it is written in.

Why the order is the bill

A model keeps nothing between requests. Each time the Chair takes a turn, the room hands it the whole Brief again from the top: who the agent is, where it is working, the tools it may call, every earlier turn, and only then the new message. A meeting of forty turns is forty complete readings, and the provider charges by the word read.

Providers soften this with a cached rate. Material identical to what the same request carried last time is read from a copy the provider kept, and the copy is cheap: according to Anthropic’s published pricing, a cached read is billed at roughly a tenth of a fresh one. The Register, the one page stating every setting of the design, records the same ratio, and this paper rests on one condition attached to it.

The condition is that a cached block counts only if everything before it is unchanged too. The copy is of a prefix, the Brief read from its first word to a marker, and one altered word before the marker makes the copy useless from that point on. So the order of the Brief is the order of its shelf life: what changes least first, what changes most last, and the message you just sent at the very end, where its arrival disturbs nothing above it.

The reference design fixes this order and marks it. Two cached blocks open the Brief: who the agent is, which changes when you rewrite its standing orders, and where it is working, your company and its open goals, which changes about daily. A budget of four cache markers is spent on them, two on those blocks, one on the tool list, and the fourth held for the conversation so far, so each turn re-reads the earlier turns cheaply and pays full price only for what is new. Three limits ride with the saving: the cache belongs to the provider, it lapses after a quiet interval, and the first reading of anything is at full price. A cached read is cheaper, never free.

The Brief, read from the top

Suppose the renewal that runs through this series, the invented client asking to keep last year’s price. Follow the Chair’s eyes down the Brief at the third turn.

Read Block What is in it How often it changes Billed
First Who the agent is Its name, its voice, the standing rules When you rewrite the standing orders From the cache, first marker
Second Where it is working Your company, the people it may address, the open goals, the date About daily, or when a goal moves From the cache, second marker
Third The tool list Every promise the room keeps, sorted by name so the order never drifts When a tool is added or withdrawn From the cache, third marker
Fourth The conversation so far The contract fetched, Finance’s floor price Grows each turn, never rewritten From the cache, fourth marker
Last The new material Your message, and whatever the Library, the room’s shelves, returned this turn Every request Fresh, at full rate

Exhibit 2. Illustrative. The Brief at the third turn of the invented renewal, read from the top: four readings from the cache, one fresh.

Four of the five readings are of material the provider has already seen, in the same order, at the cached rate. Only the last few hundred words are new. Over a long meeting the shape repeats: the cached rows grow slowly at the bottom, and the fresh row is only ever the latest message and what the Library handed over for it.

Notice what did not happen. Nobody wrote the time into the first block. Nobody put the Library’s documents into the standing orders. Nobody reordered the tool list. Each is a plausible afternoon’s work for an engineer, and each moves a change from the bottom of the Brief to the top, where it invalidates every marker below. The tenth paper takes up the documents, the eleventh the clock.

The table promises less than it seems to. A cheap re-read says nothing about how many turns a meeting takes, which belongs to the first paper. The Brief decides the unit price of a turn; the room decides how many you buy.

Fifteen is the cap on the Register, not a target. Press run and count them.

A meeting has not been run yet. Press run to watch one count, fan out to the specialists, come back, and stop.

Illustrative. One meeting of the invented renewal, counted round by round against the cap of 15.

What this arms you to ask

Ask the vendor which parts of the prompt never change between two turns of the same conversation, and which parts change on every one.

A good answer walks you down the order, names the blocks, says where the markers sit, and opens one request from the log showing how many words were read from the cache and how many fresh. That vendor has looked at their own bill. Where the answer stops short of the order, it sounds like this: one model with one prompt, assembled fresh for each request so the model always has the latest context, with caching left to the provider. Each half may be true, and together they describe a Brief rebuilt from the top every turn, with nothing for the cache to match.

Follow up with a small request: point at the first word in the prompt that changes between two consecutive turns. If it sits near the top, everything below it is read fresh, and the invoice you are holding is the consequence. If they cannot say where it is, nobody has looked.

Next paper: Why the library search never goes in the standing orders.

Asked plainly

Why does an AI agent cost more the longer a conversation runs?

A model keeps nothing between requests, so on every turn the system hands it the whole conversation again, along with its standing instructions and its tool list, and the provider charges for every word read. A forty-turn conversation is forty readings of a growing document. Providers offer a much cheaper rate for material that is identical to the previous request, but only for the unchanged run at the start; a well-ordered prompt keeps that run large.

What is prompt caching?

Prompt caching is a model provider's offer to re-read material it has already seen, in the same order, at a fraction of the fresh price. It applies to a prefix: everything from the first word up to a marker the system places. One changed word before the marker and the cached copy no longer matches, so the system must place the parts that change least at the top of the prompt and the parts that change most at the bottom.

What should be in the part of an AI agent's prompt that never changes?

Who the agent is: its name, its voice and its standing rules. Then where it is working: the company, the people it may address and the goals that are open, which change about daily. Then the list of tools it may call, in a fixed order that never drifts. Everything that arrives during a conversation belongs after all of that. The order is the whole trick, because a provider only re-reads cheaply the run of material that is identical from the first word down.

Ask the vendor

The questions this part arms, each with the good answer and the answer that arrives instead.

Learn to run the room

The certification teaches your own people to build and govern what these pages describe.