
Show me the bill for the same conversation with caching on and off.
This is a request for two numbers, and any vendor who has priced their own design already holds both. Producing them takes about ten minutes. The figures themselves are the smaller half of what you learn. The larger half is whether anybody on their side has ever opened the itemised reading cost of a single conversation and looked at how much of it was paid at full price.
Two halves of one invoice
An agent’s bill has a reading half and a writing half. Reading is the briefing handed over before every turn: standing rules, your company’s situation, the tool list, everything said so far. Writing is the replies and the specialists’ answers coming back.
Only the reading half can be cached. Providers keep a copy of material identical to what the previous request carried and charge less to read it again. According to Anthropic’s published pricing, that cheaper rate is roughly a tenth of the fresh one, and the Register, the standing page of every setting in ArkOne’s reference design for an executive agent, carries the same ratio.
Writing is billed at its own rate every time and is never cached. So the question separates a vendor along one line. A vendor who has looked knows which half of your invoice moves when the copy stops matching, and by roughly how much. A vendor who has not will answer about the total.
Forty turns, priced twice
Suppose the renewal invented for this series, taken to forty turns in one unbroken sitting. Two rooms, identical in every respect but the order of the briefing: one earns the cheap copy throughout, the other never earns it. The figures below are illustrative, modelled from the published ratio, on a briefing of a few thousand words that grows as the meeting runs, and they count only what the room reads.
| Over forty turns, modelled | Copy earned | Copy never earned |
|---|---|---|
| What is read at the fresh price | The newest message each turn | The whole briefing, forty times |
| Modelled reading cost, illustrative | USD 0.42 | USD 4.10 |
| Writing cost | Unchanged, billed at its own rate | Unchanged, billed at its own rate |
| Cost of a two-turn demonstration | Almost identical | Almost identical |
Exhibit 1. Illustrative. Forty turns in one sitting, priced twice, counting only what the room reads.
Read the last row before the second. It is why the fault reaches production. Across two exchanges the two rooms are indistinguishable, so nothing in a demonstration reveals which one you were shown.
The middle row is where a careless argument gets made, so hold the modelling steady. That multiple lands on the reading half. The writing half is unchanged, and how much the difference is worth to you depends on how much your conversations read against how much they write back. Your own figures would differ: a longer briefing widens the gap, a shorter meeting narrows it, and the ratio belongs to the provider to change.
The answer that arrives as two figures
A good answer is two figures pulled from a log, with the share of each request read from the copy beside them, and a sentence naming which half of the bill the difference sits in.
The ordinary answer arrives in one of two forms. The first is that caching is the provider’s business. True, and it answers a different question: the provider caches what matches, and the vendor’s ordering of the briefing decides whether anything matches. The second is a flat fee per seat, which makes their model costs none of your concern. That moves who carries the cost without changing how the reading is assembled, so ask what happens to the fee at renewal if the reading never caches.
Then a follow-up that needs no log at all. What is the first line of the prompt? Anything differing between two consecutive requests, a clock most often, means nothing below it can be read from the copy. The answer to your original question is then already known, and you have it before anyone opens an invoice.
Two figures, an hour of somebody’s time, before signature. The vendor who has them will produce them the same afternoon. The vendor who has to go and generate them for the first time has told you something worth more than either number. The question about what never changes is where that ordering is settled.
Asked plainly
How much cheaper is a cached read on an AI bill?
Providers publish a cached rate roughly a tenth of the fresh one, and it applies only to material identical to what the previous request carried. It covers reading, never writing. A design that earns the cheap rate through a long conversation and one that never earns it can therefore differ by close to that multiple on the reading half of the bill while doing identical work.
Does losing prompt caching multiply the whole AI bill by ten?
No, and the distinction is worth holding onto in a vendor meeting. Replies are billed at their own rate and are never cached, so that half of the bill does not move at all. The multiple falls on the reading half only. Whether that matters to you depends on how much your conversations read compared with how much they write back.
Why would a costing problem never show up in a vendor demonstration?
Because the saving only accumulates across a long conversation. A demonstration of two or three exchanges reads a short document a few times, so a design that never earns the cheap rate looks almost identical to one that does. The difference compounds with the length of the meeting and the number of meetings, which is why it is discovered on the first full invoice.
