
You were told a model’s name at the last demonstration, offered as though a name were a credential. At the next one, ask a different question. Which model answered the third specialist, and was it the one that wrote the reply?
What you see on screen is one reply in one voice, which is why the question is hard to ask. Behind it, in ArkOne’s reference design for an executive agent, sits the Switchboard: the part that chooses which model takes each call the room makes. The sentence to carry away is the title. One meeting, several models, no loyalty to any of them.
The model is chosen per call
Count the calls in one meeting. The Chair, the one voice that speaks for the room, takes a turn: a call. It consults a Seat, one of the ten specialists around the Table: another call. It searches the Library, the room’s four shelves of knowledge: a small call to a model built for matching. The first paper traced the renewal in four rounds, and those rounds were half a dozen calls. Nothing says they must all go to the same place.
The Switchboard chooses per call, by what the call needs. A Seat set to deep reasoning gets a model that thinks slowly. A question of fact whose answer sits in the Minutes, the decisions written after each meeting and read back at the next, gets a fast, inexpensive one. The Chair gets a model that writes in the company’s voice and holds a long thread. Each choice is a row in a table someone wrote and someone owns, recorded on the Register, the standing page of every setting in the design. The shape that gives your bill, the expensive model answering only the calls that need it, belongs to the ninth paper and the eleventh. This one owns who decided which model took which call, and a person did, in writing.
Two more settings ride on the Switchboard. The first is feature stripping. Models differ in what they accept: an effort setting telling a reasoning model how hard to think, a marker asking the provider to cache text. Route a call carrying such a feature to a model that does not accept it, and the design removes it before the request leaves, so the call succeeds without the feature rather than failing with it. The limit rides alongside: it succeeds with less. A stripped cache marker means a fresh read at the fresh price; a stripped effort setting means the model’s default. That is the trade, an error turned into a cost, chosen because a cost lands on an invoice where somebody can argue with it, while a failed consultation lands as a silent Seat.
The second is the search allowance: at most two web searches per meeting, each at the provider’s published rate. A room searching whenever the Chair felt unsure would run up a bill nobody budgeted, so the allowance is a number, not a mood.
One meeting, three calls, three models
Suppose the renewal again, invented for the exercise. The Chair consults Finance in the second round, Marketing and Operations together in the third, and answers in the fourth.
| Round | The call | The model chosen, and why | What the Switchboard did on the way |
|---|---|---|---|
| Second | Finance, a Seat set to deep reasoning | A slow, careful model at its low effort setting: a margin judgement with a precedent attached | Passed the effort setting through; this model accepts it |
| Third | Marketing and Operations, in parallel | A fast, inexpensive model: two questions of fact answered from the Minutes and the Library | Stripped the effort setting this model does not take; both calls succeeded |
| Fourth | The Chair, writing the reply | Chosen for the company’s voice and for holding the whole thread | Kept the cache markers; the Brief, what the Chair reads before every turn, came back at the cached price |
Exhibit 2. Illustrative. Three calls from one meeting, each answered by a different model, and what the Switchboard removed on the way.
Three calls, three models, one voice at the end, because the second paper still holds: whatever model answered Finance, the client heard the Chair. No web search was wanted, so the allowance of two stood unused. The table is also what “no loyalty” means. If the fast model in the second row doubles its price next quarter, one row changes; the Seats keep their instructions and nobody rebuilds anything.
The second row is where a move between providers goes wrong without a Switchboard. Finance’s Seat and Marketing’s share a request shape carrying an effort setting. Routed to a model that has never heard of effort settings, the request fails, and a failing consultation is a Seat that goes quiet mid-meeting, the failure the sixteenth paper called the harder one to see. Stripping makes that row a saving instead of an outage. One caution rides on all of it: a swapped model answers differently even when every call succeeds, which is why a routing change goes through the Rehearsal, the examination a change must pass first.
Fifteen is the cap on the Register, not a target. Press run and count them.
A meeting has not been run yet. Press run to watch one count, fan out to the specialists, come back, and stop.
What this arms you to ask
Ask which model answers which call, who wrote that down, and whether you may change it.
A good answer is the table above: the call types, the model for each, the reason, and who owns the list. An answer that substitutes a brand for a mechanism does the opposite, naming one model, or claiming the system picks the best model automatically. A single name means every call pays that rate, including the ones that did not need it; an automatic picker means a router nobody can show you. Ask to see the routing table and the date it last changed.
Then the half that decides whether you are buying an agent or a subscription. Move it to a different provider, and what quietly stops working? A good answer lists the features that would be removed, and what each removal costs. An answer of model-agnostic is true of almost everything and names nothing. Follow with one more: when a request carries a feature the new model does not accept, does the call fail, fail silently, or go through with the feature removed?
Next paper: Act, propose, or escalate: the one setting that decides how much your agent may do.
Asked plainly
Does an AI agent have to run on a single model?
No. The reference design chooses the model for every call the agent makes: the orchestrator's own turn, each specialist consulted, each search. Which model takes which call is a written table with an owner's name against it, and a company can change a row without rebuilding the agent.
Can I move an AI agent to a different model provider later?
In the reference design, yes, by changing the routing table rather than rebuilding the agent. What to ask about is what stops working when you do: features the new model does not accept. The design removes such a feature before the request leaves, so the call succeeds with less rather than failing, and a removed cache marker means a fresh read at the fresh price.
Why does an AI agent's bill depend on which model answers?
Because models are priced differently per unit of text, and an agent makes many calls per conversation. Sending every call to the strongest model pays the highest rate for routine work. Routing per call sends only the hard questions there. That shapes the bill; it does not fix it, because the number of calls still depends on the question.
Ask the vendor
The questions this part arms, each with the good answer and the answer that arrives instead.
