
Silent. Your commercial lead types one line into the agent: should we hold the price on the renewal. Back comes a paragraph, well judged in tone, ending in a recommendation. It is wrong, and not obviously so. Nobody queries it, because it reads exactly like the answer that would have come back had it been right. Somewhere in the build a rule had looked at that one line, seen a short question with no attachments and no tools to call, and sent it to the cheapest and fastest model on the list.
What went wrong
Two things about a question can be measured, and only one of them is easy.
The first is the shape of the request: how many words arrived, how many documents came with it, how many tools the answer needs, how long the asker will wait. All of it is countable before any work begins. The second is the weight of the answer, meaning what gets decided on the strength of it and what it costs to be wrong, and that is not in the request at all.
So a routing rule reads the shape and is quietly taken to have judged the weight. The two agree most of the day, which is what makes the arrangement durable: routine questions really are short, and answering them on something light really does save money. Then a one-line question arrives carrying a pricing decision, and the saving is taken on the one call of the week where it should not have been. A lighter model does not announce that it was out of its depth, having no means of knowing, and fluency is not connected to depth, so the reader gets no signal. The same gap elsewhere produces a weak match quoted with full confidence: there the missing judgement is about a document, here about a model.
In the Cabinet, ArkOne’s reference design for an executive agent, this belongs to the Switchboard, the part deciding which model takes which call. The model is chosen per call rather than per system, so one meeting can draw on several and owes none of them loyalty. A feature the chosen model does not support is stripped out before the request leaves, so routing a call elsewhere can never turn a working call into an error, which is what otherwise drives a build to give up and standardise on one model for everything. And the four deep-reasoning Seats, the specialists whose job is to think rather than fetch, sit at the low effort setting, because higher settings cost roughly three times the time for no gain the standing examination could measure, a modelled figure rather than a billed one. That is the shape worth copying: a cheaper setting chosen because it was measured, not because it was cheap. All sit on the Register, and the eighteenth paper sets out the routing.
The record
Suppose a specialist software reseller of seven hundred and fifty people, whose agent answers commercial questions for the sales desk. The reseller, the two questions and the model tiers are invented for this exercise.
| The question, as typed | What the rule saw | Where it went, and what came back |
|---|---|---|
| ”Does the Leeds office have Saturday cover” | Short, no documents, no tools, reply wanted now | The light model. Correct, instant, and the right place for it |
| ”Should we hold the price on the Harrogate licence renewal” | Short, no documents, no tools, reply wanted now | The light model. A confident recommendation, reasoned from nothing but the wording |
Exhibit 1. Illustrative. Two questions of the same shape, one trivial and one consequential, sent by the same rule to the same model.
The rows are indistinguishable in every column the rule can read, and that is the finding. No bug is locatable: the rule inspected the request accurately and classified it correctly on the terms it was given.
What separates them sits downstream of the agent. The first answer is checked within a minute, against the office rota, by the person who received it. The second is checked by nobody, because there is nothing to check it against, and it enters a negotiation instead. The cost is paid in a margin months later, when nobody is asking which model answered a question in the spring.
Notice what a cost report shows for that week. Both calls appear as savings, the light model handled a high share of the volume, and the average cost per question fell. A routing rule measured on cost will always report success, because cost is the half of the trade it can see.
What it means for you
Decide the model where the stakes are already known, which is at the specialist rather than at the request. Nothing in a short question can tell software what the question is for, so asking the software to route by importance is asking it for a judgement it has no material to make. Asking who is answering is a question somebody has already settled on purpose.
The specialist answering office hours and the specialist reasoning about pricing are different seats in the room, each carrying its own model and its own standing. The routing question then stops being how big this message looks. Each specialist also carries one authority setting, act alone, propose and wait, or escalate to a person, and the one answering pricing questions is the obvious candidate for the second.
Name one decision your company took this year that you would hate to have got wrong, and ask which model answered the question behind it. Then ask to see the rule that chose that model and what the rule looks at. If the answer is the length of the question, how quickly a reply was wanted, or how many tools it needed, you have been told the price of your agent and nothing about its judgement. A build that measured its way to a cheaper setting, as the eighth paper works through, can show you what it measured against.
Then ask what happens when a model in the list does not support something a call needs. A build with an answer has thought about routing. A build without one finds out in production, and repairs it by sending everything back to a single model, which is where this note began.
Asked plainly
How does an AI agent decide which model answers a question?
A rule somewhere in the build picks one, and what that rule reads is the important part. Most routing rules read properties of the request: how long it is, how many tools it needs, how fast a reply is wanted. None of those describe how much the answer matters. A question can be eight words long and decide a contract, and a rule reading length cannot see the difference.
Can a cheaper AI model give a confidently wrong answer?
Yes, and that is the difficulty. A smaller model does not know it is out of its depth and has no way to report it, so a thin answer arrives in the same fluent, measured prose as a thorough one. The reader has no signal to work from, which is why the choice of model has to be made deliberately rather than discovered afterwards from the results.
Is it cheaper to run an AI agent on one model for everything?
It is simpler to build and it is rarely cheaper to own. One model for everything means either paying a heavyweight rate to answer trivia, or answering consequential questions with something too light for them. Choosing per call costs a little engineering once and lets the same conversation draw on a cheap model for routine steps and a stronger one where the decision sits.
