ARKONE
← All questions
Questions for the Vendor, No. 13The Brief · the standing instructions4 minute read

What is in the part of the prompt that never changes, and what is in the part that does?

The briefing is re-read every turn, and its order, not its length, is what decides the reading half of your bill.

The Brief, seen from above on the Cabinet floor plan.
Exhibit 1. The Brief, at its place in the room.

What is in the part of the prompt that never changes, and what is in the part that does?

The invoice is the reason you ask. It arrived larger than the demonstration suggested, the price per word looked modest, and nothing in the conversations you read accounts for the gap. The gap sits in a document nobody showed you: the briefing the agent is handed before every single turn, containing its standing rules, your company’s situation, its tools and everything said so far.

Shelf life, top to bottom

A model keeps nothing between requests. Every turn, the briefing goes over again from the first word, so a long conversation is many complete readings of a document that grows as it runs.

Providers reduce this with a cached rate for material identical to what the previous request carried. According to Anthropic’s published pricing, that rate is roughly a tenth of the fresh one, and the same ratio sits on the Register, the standing page of every setting in ArkOne’s reference design for an executive agent.

One condition attaches, and the whole question rests on it. The cheap copy covers an unbroken run from the very first word to a marker. Change one word before the marker and everything below it is fresh again. So order is shelf life: standing rules first, working context next, the tool list in a fixed order, the newest message last. The reference design opens with two cached blocks and spends a budget of four markers, holding one for the conversation itself.

The same briefing, ordered two ways

Suppose one meeting of the invented renewal that runs through this work, at its tenth turn, with a briefing of a few thousand words. Two rooms hold identical material. One puts the standing rules first; the other opens with a line carrying the time to the second. Both readings are modelled from the ratio above.

Block, in reading order Changes how often Rules first Clock first
Who the agent is When you rewrite the rules Cheap copy Fresh
Your company and its goals About daily Cheap copy Fresh
The tool list, sorted by name When a tool is added Cheap copy Fresh
The conversation so far Grows each turn Cheap copy Fresh
The newest message Every request Fresh Fresh

Exhibit 1. Illustrative. One briefing, two orderings, and how much of each is re-read at the fresh price.

One column pays the fresh price for the last few hundred words. The other pays it for the whole document, ten times over by the tenth turn, for work that is otherwise identical: same tools, same specialists, same reply.

The multiplier applies to the reading half only. What the room writes back is billed at its own rate and is never cached, so that half does not move. A two-turn demonstration loses almost nothing, which is why this is found on an invoice rather than in a meeting.

The answer that walks down the order

A good answer walks down the order block by block, says where the markers sit, and opens one request from the log showing how many words were read cheaply and how many fresh. That vendor has read their own bill.

The ordinary answer sounds careful and describes the opposite. The prompt is assembled fresh for each request so the model always has the latest context, and caching is handled by the provider anyway. Both halves can be true at once, and together they describe a briefing rebuilt from the top every turn, with nothing for a cheap copy to match. The provider caches what matches; the vendor’s ordering decides whether anything does.

Follow up with a request that needs no log and no engineer. Point at the first word in the prompt that differs between two consecutive turns of the same conversation. If it sits near the top, everything below it is being read at full price on every turn, and your invoice is the arithmetic. A timestamp is the common case, but a request identifier, an unread-message count or a tool list assembled in whatever order the tools happened to load all do the same work. If nobody can say what the first line is, nobody on their side has looked.

Order is not the only thing that decides a bill, and it is worth saying which part it owns. The briefing decides what one turn costs to read. How many turns a meeting takes is a different setting with a different owner. The bill shown twice is where this question is settled with figures.

Asked plainly

Why does an AI agent re-read its whole prompt every turn?

Because the model keeps nothing between requests. Each turn, the system hands over the standing instructions, the working context, the list of tools and every earlier exchange again from the top, then adds the new message. A long conversation is therefore many complete readings of a growing document, and you are billed for the reading as well as for what comes back.

What is prompt caching and why does the order of the prompt matter?

Providers keep a copy of material identical to what the previous request carried and charge less to re-read it. The copy covers an unbroken run from the very first word, so one altered word near the top makes everything below it fresh again. That is why designs put what changes least at the top and the newest message at the very end.

What should be at the top of an AI agent's prompt, and what at the bottom?

At the top, whatever changes least: who the agent is and the standing rules it works under. Then the working context, which changes about daily. Then the tool list, in a fixed order. At the very bottom, the newest message and anything just fetched, because material arriving at the end disturbs nothing above it.

Talk it through before you decide

A discovery call, no deck: your situation, the parts of the room it touches, and what you would need to decide first.