
Somewhere in the vendor’s demonstration is a slide showing the prompt. At the top sits a paragraph about who the agent is, and beneath it, pasted in, the three documents found for the question on screen. The vendor calls this giving the model your context. On the invoice it is the most expensive place in the request to put a search result.
What you are looking at on that slide is a decision about placement. The rule governing it in the Cabinet, ArkOne’s reference design for an executive agent, takes one sentence: anything the room finds during a meeting arrives at the end of the Brief, the document the Chair reads before every turn, and never in the standing part the Chair was already carrying. This paper is about what that rule saves, and what it costs to break.
Fresh material is the most changeable thing in the room
The Library is the room’s four shelves of knowledge, kept apart: house knowledge, the client’s documents, recorded failures, recent research. Asked a question by the Chair, the one voice that runs the room, it searches the shelf it was pointed at and returns the passages matching well enough. Well enough is a setting on the Register, the page stating every setting of the design: a match beyond the recorded similarity distance is discarded, because a weak match is worse than none. The thirteenth paper is about that dial; this one is about where the passages that pass it go.
The ninth paper set out the order of the Brief: the provider re-reads unchanged material at a fraction of the fresh rate, but only the run from the first word down to a marker, so what changes least stands first and what changes most stands last. Two cached blocks open the Brief, who the agent is and where it is working, and the conversation so far follows behind a marker of its own.
Now ask what a search result is, in those terms. It exists because of one question at one turn and will differ at the next; nothing in the request changes more often. Written into the first block, it makes that block differ from the previous request, and since every marker below depends on the run above matching, the whole Brief is read fresh, this turn and every turn the search runs. Appended after the conversation, the same passages disturb nothing: the standing blocks match, the earlier turns match, and only the new words are fresh.
There is a second reason, about authority rather than money. The standing part is where the room tells the Chair who it is and what it must always do, and material placed there is read with that weight. A contract fetched for one pricing question has no business beside the agent’s name, and a fetched document carrying a line that looks like an instruction is the strongest case of all. Placed at the end, the same document is evidence: read, weighed, gone when the meeting ends.
The limit, plainly: appended material still lengthens the conversation, and a meeting that searches at every turn grows a tail. That tail is re-read at the cached rate, which is cheaper and never free, and discarded with the meeting.
The same two documents, in two places
Suppose the renewal that runs through this series, the invented client asking to keep last year’s price. At the first turn the Chair asks the Library for the contract and last year’s pricing, and two documents come back. Follow them through two rooms that differ only in where the documents are put.
| Moment | Placed after the conversation so far | Placed in the standing part |
|---|---|---|
| The documents arrive | Appended behind the fourth marker, after the Chair’s question | Written into the first block, beside who the agent is |
| The request leaves | Standing blocks, tool list and earlier turns from the cache; the documents and question fresh | The first block no longer matches, so nothing below it matches; the whole Brief fresh |
| Later turns, to the answer | Every earlier turn, documents included, from the cache; only the newest reply is fresh | Each turn a full fresh reading of a growing Brief |
| Tomorrow, a different client | The documents left with the meeting; the standing blocks are as they were | Yesterday’s pricing sits in the standing orders; another full fresh read, the wrong client’s contract beside the agent’s identity |
Exhibit 2. Illustrative. The same two documents from the Library, placed after the conversation in one room and in the standing part in the other.
The two rooms found the same documents and reached the same answer. One paid the fresh rate for two documents and a question; the other paid it for the entire Brief, four times in one meeting, on a decision made once and repeated silently on every request afterwards.
The last row is the one to keep. In the second room the result stayed, in a place that grants it standing, until the next search wrote over it, so yesterday’s client sat beside the agent’s own name in a Brief read for a different client. The reference design cannot make that mistake: no path runs from a Library search to the standing part.
What this arms you to ask
Ask the vendor where a document the agent has just found is placed.
A good answer puts the found material after the conversation and shows a request in which the standing blocks came from the cache while the documents were read fresh. The answer that should worry you names no mechanism: the relevant documents are given to the model so it always has your context, or the prompt is dynamic and personalised for every request. Both describe a first block that differs on every turn, without saying so.
Follow up by asking for one request where the agent used a document, and where in that request it appeared: before the conversation or after it. Then ask what became of it on the next request. If it was replaced, the standing part is rewritten on every turn, and you pay for the whole Brief each time.
Next paper: The one line that makes a bill ten times larger.
Asked plainly
Where should an AI agent put the documents it retrieves?
After everything that does not change. A retrieved document exists because of one question at one turn and will be different at the next, so it is the most changeable material in the request. Placed at the end, after the standing instructions, the tool list and the earlier turns, it disturbs nothing and is read fresh only once. Placed at the top, it changes the beginning of the prompt on every turn and forces the provider to re-read the whole thing at the full rate.
What goes wrong when an AI agent pastes a found document into its standing instructions?
Two things. Cost: the cheaper cached rate applies only to the unchanged run from the first word down, so a document that differs on every turn, sitting at the top, means nothing below it can be read cheaply. Authority: the standing instructions are where the agent is told who it is and what it must always do, and material placed there is read with that weight. A document fetched for one question is evidence to be weighed, and if it happens to carry a line that reads like an order, the standing instructions are the one place it should never sit.
Does a search make an AI agent conversation more expensive?
Yes, and the amount depends on where the result lands. Placed at the end of the prompt, the result is read fresh once and then re-read at the cached rate for the rest of the conversation, so the added cost is the result itself. Placed in the standing part, it invalidates the cache for the whole prompt on every turn the search runs, and the added cost is the entire prompt, repeatedly.
Ask the vendor
The questions this part arms, each with the good answer and the answer that arrives instead.
