
Silent. Somebody asked what competitors were charging, and the answer that came back was excellent. Four paragraphs, a figure, a comparison to the company’s own list price, a recommendation. It was circulated. Nobody queried it, because there was nothing in it to query. Two of the four claims were the company’s own approved numbers. One was a figure somebody had posted on a forum that morning, and the answer gave you no way to tell which was which.
What went wrong
Material fetched this morning and material your board approved last quarter do not deserve the same confidence, and by the time an answer is written it is usually too late to tell them apart.
The Library, the reading room an agent consults when it needs something it was not handed, keeps four shelves in ArkOne’s reference design for an executive agent, the Cabinet: what the design knows about running an executive meeting, your own documents, failures somebody recorded after they happened, and anything fetched from outside during the meeting. The Register sets the count at four, and the twelfth paper argues why they are kept apart.
Separate shelves are half the job. A passage comes back from the Library carrying the shelf it sat on, so the Chair, the one voice that speaks for the whole room, reads a web page as a web page. The other half is whether that label survives into the answer.
It usually does not, and the reason is mechanical rather than careless. The model receives a set of passages and writes prose. Prose has one voice. Fluency is set by the model and owes nothing to sourcing, so a sentence built on an anonymous comment reads exactly as well as a sentence built on an audited figure. Unless something requires each claim to name its origin, the reply flattens four kinds of authority into one confident register, and the flattening is invisible by construction.
The record
Suppose a competitor pricing question, with the company, the figures and the forum post all invented for the exercise. Four claims left the room in one answer.
| The claim, as the reader saw it | Where it actually came from | What it was worth |
|---|---|---|
| Our list price is a stated figure per seat | The company’s own price list, approved | Solid, and checkable in a minute |
| Our margin floor sits at a stated percentage | An approved finance paper | Solid |
| The nearest competitor charges a stated figure less | A comment on a public forum, fetched that morning | One person’s assertion, undated, unverified |
| We should therefore consider matching | The model, reasoning across the three above | Only as good as the weakest input |
Exhibit 1. Illustrative. Four claims in one answer, and where each of them actually came from.
The first two rows are the reason the fourth was believed. A recommendation sitting on top of two verifiable facts and one anonymous comment inherits the credibility of the facts, because the reader checks what is checkable, finds it correct, and extends the benefit to the rest.
Nothing in the answer marked the third row as different. It was not hedged, not attributed, not dated. Had it said that a public forum comment from that morning put the competitor lower, the reader would have known exactly what to do with it, which is to ask somebody to confirm it before repricing anything.
This is why the note is tagged silent. There is no incident here. The answer arrived, read well, was circulated and acted on, and the failure only becomes visible if the competitor’s real price is discovered later by other means. Most of the time it is not, and the organisation quietly holds a belief it never examined.
What it means for you
Provenance is a property of the answer rather than of the archive. A company can store its documents beautifully, in properly separated stores, and still receive replies in which every claim sounds equally authoritative. The separation is what makes labelling possible; it does not perform it.
The practical consequence is about what you may safely delegate. An answer whose claims each name their origin can be checked in the time it takes to read the weakest one. An answer that does not can only be trusted wholesale or distrusted wholesale, and neither is a workable basis for a decision, so in practice the good writing wins and it gets trusted.
Ask whoever built your agent to show you a reply where a claim came from a web page fetched during the run, and ask where in that reply the reader is told. Then ask the sibling question about the shelf that decides relevance, which an earlier note takes up: what is discarded for being too weak a match to quote at all.
Asked plainly
How do I know where an AI agent got a number from?
Only if the answer tells you. Every claim in an agent's reply came from somewhere: a document your company approved, a page fetched from the web minutes earlier, or the model's own general reading. Those three deserve very different amounts of trust, and unless the reply labels each claim with its origin, they arrive in one confident voice and you have no way to sort them.
Can an AI agent mix up company documents with things it found online?
It can, and the mixing usually happens before the agent ever sees the material. If everything the agent reads is stored in one pile, a search returns passages with no note of which pile each came from. Keeping approved documents, recorded failures and material fetched from the web in separate stores is what makes the difference visible later.
Why does a wrong figure from an AI agent sound so convincing?
Because fluency is not connected to sourcing. The writing quality of a sentence is set by the model, and it is the same whether the underlying figure came from an audited board paper or an anonymous comment. A weakly sourced claim therefore arrives in the same measured tone as a strong one, which is precisely why it survives a read.
