ARKONE
← All questions
Questions for the Vendor, No. 15The Switchboard · model routing4 minute read

Which model answers which kind of question, and who decided?

One conversation is several calls, and a design either has a written table saying where each one goes or it does not.

The Switchboard, seen from above on the Cabinet floor plan.
Exhibit 1. The Switchboard, at its place in the room.

Which model answers which kind of question, and who decided?

At the last demonstration somebody named a model, in the tone people use for a credential. You are asking the question that the name does not answer. Behind one reply on your screen sit several separate calls, and the model that wrote the polished paragraph you were shown is not necessarily the one that answered the third specialist, priced the margin, or searched your documents.

The unit is the call

Count the calls rather than the conversations. The main voice takes a turn, which is one call. It consults a specialist, which is another. It consults two more at once, which is two. It searches the stored documents, which is a small call to a model built for matching. Then it writes the reply.

The reference design chooses per call, by what the call needs, and that choice sits in a table somebody wrote and somebody owns. The Register, the standing page of every setting in ArkOne’s reference design for an executive agent, records the granularity as per call, with web searches rationed at two per meeting so a room that feels unsure cannot spend without a limit.

One protection rides alongside. Where a chosen model does not accept an extra the call carries, the extra is removed before sending, so the call succeeds at a slightly higher cost rather than failing silently.

Six calls, two ways of routing them

Suppose one meeting of six calls, working the renewal invented for this series. The middle column routes each call by what it needs. The right-hand column sends all six to one capable model, which is what a single named model in a demonstration means in practice.

The call Routed by need Routed to one model
The main voice, opening the meeting Model chosen for voice and long threads Same model
Finance, weighing a margin A slow, careful model Same model
Marketing, a fact from the record A fast, inexpensive model Premium rate
Operations, a fact from the record A fast, inexpensive model Premium rate
Searching your documents A small matching model Premium rate
The main voice, writing the reply Model chosen for voice Same model

Exhibit 1. Illustrative. Six calls in one meeting, and what each costs when routed by need against routed to one model.

Half the calls in this meeting are lookups. In the right-hand column all three are billed at the rate of the most capable model in the building, for questions where a slow, careful answer was never wanted. Multiply by the meetings in a month and the difference is the difference between two invoices for identical work.

The table is also what independence looks like. If the fast model doubles its price next quarter, one row changes and nobody rebuilds anything. Where no table exists, there is no row to change.

The answer that is a table with a date on it

A good answer is that table: the call types, the model for each, the reason, the person who owns the list, and the date it last changed. Ask for the date. A list nobody has revised since a price moved is a list nobody owns.

The ordinary answer replaces a mechanism with a brand. One model name, offered as the whole architecture. Or the claim that the system automatically picks the best model for the task. A single name means every call in the meeting pays that rate, including the three that did not need it. An automatic picker means a router nobody can show you, and a routing decision you cannot inspect is one you cannot change when your bill changes.

Then ask the half that decides whether you are buying a design or a subscription. Move it to a different provider next quarter, and what quietly stops working? Model-agnostic is true of almost everything and names nothing. Press for the list of extras that would be removed and what each removal costs. Close with the sharpest version: when a call carries something the new model does not accept, does it fail, fail silently, or go through with the extra removed?

A change of model changes the answers as well as the price, so a vendor who routes deliberately should also be able to say how a swap is tested before it reaches you. The bill shown twice is the other half of the same invoice.

Asked plainly

Does one AI agent conversation use one model or several?

One reply on screen can rest on several separate calls. The main voice takes a turn, each specialist consulted is another call, and searching stored documents is another again. Nothing requires them all to go to the same model, and a design that routes each call by what it needs pays the expensive rate only where a slow, careful answer was actually wanted.

Who decides which model answers which question inside an AI agent?

Either a person, in a written table naming the call types and the model chosen for each, or nobody in particular. A vendor able to show you that table and the date it last changed has made the decision deliberately and can change it when a price moves. A vendor describing an automatic picker is describing something you cannot inspect or argue with.

Can I move an AI agent to a different model provider later?

Sometimes, and the useful question is what quietly stops working when you do. Models differ in the extras they accept, such as an instruction about how hard to think or a marker asking the provider to keep a cheap copy of the text. A design that removes an unsupported extra before sending keeps the call working at a slightly higher cost, rather than failing it.

Talk it through before you decide

A discovery call, no deck: your situation, the parts of the room it touches, and what you would need to decide first.