ARKONE
← All papers
Cabinet Paper No. 11The Brief · the standing instructions6 minute read

The one line that makes a bill ten times larger

A line that changes on every request, placed first, makes everything after it fresh, and what the agent reads costs roughly ten times more.

The Brief, seen from above on the Cabinet floor plan.
Exhibit 1. The Brief, at its place in the room.

An engineer adds one sensible line to the top of the agent’s prompt. It tells the model the date and time, to the second, so a reply about a deadline is never a day out. Nothing breaks. The answers are as good as they were. The following month your invoice is far larger than expected, and nobody can point to what changed.

This is the third paper on the Brief, the document an agent is handed before every turn, read by the Chair, the one voice that runs the room. Both belong to the Cabinet, ArkOne’s reference design for an executive agent. The two papers before set out why the order of the Brief decides its cost; this one shows how little it takes to lose the saving. One line that differs on every request, placed first.

A clock face at the top of the page

The mechanism is the one the ninth paper described. A model keeps nothing between requests, so the room hands it the whole Brief every turn. The provider re-reads the run identical to the previous request, from the first word to a marker, at a cached rate; according to Anthropic’s published pricing, that rate is roughly a tenth of the fresh one, and the Register, the page stating every setting of the design, carries the same ratio. Everything below the first changed word is read fresh.

Put a clock at the top and the first changed word is the first word. The seconds have moved since the last request, so the first block no longer matches, and since every marker below depends on the run above matching, none matches either. The standing orders, the description of your company, the tool list and every earlier turn are read fresh, on every turn, for the life of the agent. The cache is never earned, and the request looks from the outside exactly as before.

The arithmetic follows from the ratio, and it applies to one half of the bill. When nearly all of what a request reads could have cost a tenth and instead costs full price, the reading half is multiplied by roughly ten. What the room writes back, every reply and every specialist’s answer, is billed at its own rate and never cached, so that half does not move. A two-turn meeting loses almost nothing, which is why the mistake survives every demonstration and is found on an invoice.

The reference design keeps the clock and places it where it can do no harm. Today’s date belongs in the second block, where it is working, which already changes about daily; a date that moves once a day costs one fresh reading of the blocks below, once a day. The time to the second, if a question needs it, belongs in the newest message at the end, where its arrival disturbs nothing above. The rule is general: a thing that changes on every request goes last.

The mistake wears other costumes. A tool list assembled in whatever order the tools loaded is sorted by name before every request, for the same reason. A request identifier, a random greeting, an unread count: each is a clock face by another name.

Two bills for the same forty turns

Suppose the invented renewal from the first paper, run to forty turns in one sitting, in two rooms differing by one line. The first room’s Brief opens with who the agent is; the second opens with the date and time to the second and is otherwise identical. Because the cache lapses after a quiet interval, the comparison has to sit inside one unbroken meeting, not across a working day.

Turn First line never changes First line is the clock
First Everything fresh once, and the copy is written Everything fresh, and the copy will never match again
Second Standing blocks, tool list and the first turn from the copy; only the new message fresh The seconds moved, so nothing matches and everything is fresh
Fortieth Forty new messages fresh in total, the rest cached Forty complete readings of a growing Brief, all fresh
Modelled reading bill for the forty turns, illustrative USD 0.42 USD 4.10

Exhibit 2. Illustrative. Forty turns of the invented renewal in one sitting, in two rooms; one opens its Brief with a clock, and its modelled reading bill is roughly ten times larger.

Those figures are illustrative, from the Register’s worked illustration of the ten-times bill: forty turns in one sitting, a Brief of a few thousand words growing as the meeting runs, priced from the cached-read ratio, counting only what the room reads, since replies are billed separately and never cached.

The two rooms did identical work: same documents, same specialists, same reply through the Door, the rules that decide what may leave the room. The one that knew the time to the second paid roughly ten times as much for its reading, and would have known the time just as well from the end of the Brief.

Your figures would differ. A longer Brief widens the gap, a shorter meeting narrows it, and the ratio is the provider’s to change. The direction does not move, and no other setting recovers what that first line spends.

What this arms you to ask

Ask the vendor to show the bill for the same conversation with caching on and off.

A good answer is two figures from a log, with the share of each request read from the cache beside them. That vendor has measured what this paper is about. Two answers should send you back to the invoice. The first is that caching is handled by the provider, or that a flat fee per seat makes their model costs none of your concern. The first is true and does not answer you: the provider caches what matches, and the vendor’s prompt decides whether anything does. The second moves who carries the cost without changing how the reading is assembled. Ask what happens to that fee at renewal if the reading never caches.

Follow up with a question needing no log: what is the first line of the prompt? If it carries the time to the second, or anything else differing between two consecutive requests, you have your answer without seeing the bill. If the vendor cannot say what the first line is, nobody on their side has looked.

Next paper: Four shelves, kept apart on purpose.

Asked plainly

What breaks prompt caching in an AI agent?

Any word near the top of the prompt that differs between one request and the next. A timestamp to the second is the common case; a request identifier, a random greeting, an unread-message count or a tool list whose order drifts does the same. The provider's cheaper rate applies only to the unchanged run from the first word down, so one moving line placed first means nothing after it can be read from the cache.

How much more does an AI agent cost when caching never hits?

Providers publish a cached-read rate that is roughly a tenth of the fresh rate, and it applies only to what the agent reads, never to what it writes back. In a long conversation nearly all the reading could be at the cheaper rate, so losing the cache multiplies the reading half of the bill by roughly ten while the replies are billed as before. A two-turn conversation loses almost nothing, which is why the fault survives a demonstration and is found on an invoice.

Should an AI agent know the current date and time?

Yes, and where the date sits decides what it costs. A date that changes once a day belongs in the part of the prompt that already changes about daily, where it costs one fresh reading per day. The time to the second, if a question ever needs it, belongs in the newest message at the very end, where its arrival changes nothing above it. What must never happen is a clock in the first line.

Ask the vendor

The questions this part arms, each with the good answer and the answer that arrives instead.

Learn to run the room

The certification teaches your own people to build and govern what these pages describe.