Read in parallel, decide in one place
Do not build multi-agent systems, said one lab in 2025; multi-agent systems are working, said the same lab in 2026. Both are right once you separate reading from writing. In care the implicit decisions behind an answer are specific, and the line between reading and deciding has to be drawn in the harness, not left to the model.
Written by
Founder and CEO
Essay3 min read
In June 2025 Walden Yan at Cognition published Don’t Build Multi-Agents. His point: every action carries implicit decisions, and when several agents work in parallel each makes them in isolation, so the pieces do not fit. Last month the same author wrote Multi-Agents: What’s Actually Working, and it reads like a partial retreat. It is not. The two pieces describe one rule seen from two sides, and I think it is the most useful rule I have read this year for building anything where answers matter. Cognition sells a coding agent, so its interest is in the subject.
One writer, many contributors
Many can read. One writes. Parallel readers, each sent off to find something, lose nothing if their results are brought together in one place. Parallel writers each decide something, and the decisions collide. Anthropic’s Prithvi Rajasekaran adds a second half: whoever checks a piece of work should be a separate, sceptical party, because a system grading its own work is mild towards it. Dex Horthy’s 12-Factor Agents puts the same idea in a sentence I will not copy: a request for a tool is a proposal, and ordinary code decides whether it goes ahead.
What the implicit decisions are in care
In software the implicit decisions are style and structure. In care they are concrete. Which client is this about. Which version of the protocol applies today. Which source the organisation’s working agreements designate for this subject. Which region’s agreement counts. Four agents each answering a quarter of that, without seeing each other’s designation, will produce an answer that sounds coherent and rests on a mixture.
So the design I want is this. Retrieval can fan out, because reading changes nothing. The answer is composed in one place that sees every designation. The professional is the last link in that single thread. A verifier starts without the history of the drafting, but not without the sources: it checks against the version in force, not against what the draft believed.
The line is in the harness
Yan admits that weaker models do not know their own limits, and expects training to close the gap. I do not wait for that. For medication, doses and anything that crosses the line Inora is not to cross, escalation should be a fixed rule in the harness, so that it can be audited and so that it is specific to the organisation. Two copies of the same model probably share the same blind spots in Dutch care regulation, so a second reader of the same make is a weaker check than it looks.
One more thing I would not do: give a consulting model a full copy of the context. Inside a sovereign boundary that model has to run within it, or it receives only a minimised question.
Not a dogma
I will not repeat “never use several agents”. The real boundary is between reading and writing, and between proposing and deciding. I would turn Horthy’s direction around, though. In his picture the human approval is something the agent can call. For us it is where the decision lives, and the agent’s work is what is put in front of it.