
How many specialists can it consult at once, and what stops it consulting all of them every time?
You get one clean shot at this, at the moment they invite you to type your own question. Give the agent something genuinely wide. Not a lookup. Something like: our largest competitor has cut prices in our main market, what should we do this quarter. A question that broad should pull several specialists in. Ask both halves while it runs, then watch three screens rather than the reply.
A ceiling, and the thing that is not a brake
An agent with specialists puts one question to several of them in the same round and waits for them together, so the round takes as long as the slowest reply rather than all of them added up. The limit on how many may be named in one request is the fan-out cap.
In the Cabinet, ArkOne’s reference design for an executive agent, that cap is set at the size of the roster, which is ten specialists, listed alphabetically so that no order is implied: Board, Counsel, Finance, Marketing, Operations, People, Product, Strategy, Talent and Triage. Ten on the list, ten askable at once. A specialist over the cap is handed an explicit skip message rather than dropped in silence, so the room always knows who was not asked. Both values sit on the Register, the standing page of every setting in the design.
Which leaves the second half a plain answer. A ceiling at the roster permits asking everybody every time. What restrains it is judgement about which views the question needs, and the fact that each reply is paid for and then read back by the voice that has to speak.
Three screens while it runs
Suppose you have put the competitor question in, in a demonstration invented for the exercise. Ignore the answer for the moment and look at these instead.
| Where to look | The reading you want | Why this screen |
|---|---|---|
| Timestamps in the transcript | Several specialist requests carrying the same second, then replies landing at different seconds | Requests together and replies apart is what working in parallel looks like from outside |
| Which specialists were named | Three or four, chosen for this question, with the rest of the roster untouched | Choosing is the judgement you are paying for; asking everybody is not |
| The usage or cost panel, if they will open it | A count of calls that matches the specialists you just counted | The bill is the only thing that cannot be rehearsed |
Exhibit 1. Illustrative. One question put to the agent live, and the three screens that show whether specialists really answered together.
The first row is the one nobody can stage. Requests sharing a timestamp and replies arriving apart is the signature of several calls in flight together, and no amount of narration produces it.
The second row answers the second half of the question as behaviour rather than as a sentence. A room that asks all ten of a pricing question has no view about what a pricing question needs, and you will meet that again in your bill every week.
The number, the shrug, and the follow-up
A good answer has two parts and comes quickly. A number, where it is set, and what a specialist over the ceiling receives. Then the honest half: the ceiling permits everybody, and what stops the agent asking everybody is judgement you can inspect afterwards in the transcript.
The reply that sounds better and says less is that the agent draws on whatever expertise the question needs. That may describe something excellent. It may also describe one model with one long set of instructions weighing the finance angle and the legal angle inside a single reply, in which case there is nobody to consult. The follow-up cannot be talked around: show me two specialist replies that arrived in the same round, with their timestamps.
Then the question about restraint, put as arithmetic rather than architecture. If it consulted every specialist on every request, what would a week cost, and what prevents that today? A vendor who has watched a bill has a figure and a policy. A vendor who has not will offer that it is configurable, which describes a field on a settings page.
What the wide question buys you
Ask this once and three claims get tested together: that the specialists exist, that they work together rather than in a queue, and that somebody decided which ones a given question deserves. How a room asks many and hears them at once is worked through in the Cabinet paper on parallel consultation. What each specialist costs per call is the question about which model answers.
Asked plainly
How many specialists can an AI agent consult at the same time?
That depends on a ceiling somebody set, usually called a fan-out limit. A sensible default is the size of the roster, so every specialist on the list may be asked in one round if the question genuinely needs them. What matters more than the number is that a specialist left out of a round is told so rather than silently dropped.
Are AI specialists real separate agents or just headings in one answer?
Both arrangements exist and they look identical in a polished reply. A separate specialist is its own call to a model with its own instructions, so several requests carry the same timestamp and each returns its own text. Headings in one document arrive as one call, and no transcript will show two replies arriving together.
What stops an AI agent from asking every specialist every time?
Nothing in the ceiling itself, which is why the question has two halves. Every consultation is a reply that must be paid for and read back, so the restraint is the orchestrator's judgement about which views the question needs, plus whatever limit exists on how many rounds a single job may take.
