Knowing When Not To Add Another Agent Becomes The Core Skill In Building Reliable AI
Cognition co-founder Walden Yan says the reflex to add another agent is what makes AI systems fragile, and reliability comes from controlling context.

If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.

Adding another agent is the instinctive fix for an AI product that keeps getting things wrong. Split the task, assign a specialist to each piece, and let a coordinator merge the results. The system that comes out the other side is often less reliable than the single agent it replaced.
Cognition, the company behind the Devin coding agent, makes that case in a post titled "Don't Build Multi-Agents." Co-founder Walden Yan described the swarm-of-subagents design as very fragile and puts reliability somewhere less exciting than architecture. "At the core of reliability is Context Engineering," he said.
Where reliability comes from
Two products built on the same frontier model start with the same intelligence. What separates them is reliability, and reliability comes from controlling context.
Subagents working in parallel can't see each other's work. Each one guesses at what the others are doing. Yan's example splits a Flappy Bird clone into two jobs. One subagent builds a background that looks like Super Mario, the other builds a bird that moves nothing like the game, and a final agent has to combine them. "Actions carry implicit decisions, and conflicting decisions carry bad results," Yan explained.
The fix costs more than handing each subagent the original task. What shapes an agent's judgment lives in the turns and tool calls that came before it, so a summary leaves out the part that matters. "Share context, and share full agent traces, not just individual messages," he noted.
Adding parts only when they earn it
Building a reliable agent means turning down complexity. Anthropic's guidance for agent builders says to find the simplest system that works and to add "complexity only when it demonstrably improves outcomes."
Its one case for splitting the work is a narrow one. A separate model that screens inputs for what the main agent shouldn't handle "tends to perform better" than asking a single model to answer and police itself, Anthropic noted. The second part earns its place by preventing a named failure.
When the setting is a live buyer call
A live sales call is an unforgiving place to run an autonomous system. The conversation runs long and unscripted, the buyer interrupts and doubles back, and the system has to stay coherent the whole way. There's no clean seam to split the call along, and no way to reconcile two subagents that heard different halves of it.
1mind builds for that setting. Its Ride-Along product puts a photorealistic, voice-enabled teammate onto live Zoom, Teams, and Meet calls as a named participant that answers the buyer out loud. Frontier models from OpenAI and Google Gemini supply the intelligence. Keeping that intelligence coherent as a single participant across an hour of live conversation is a context-engineering problem, and it's the one 1mind builds against.
Trust in an autonomous teammate comes from coherence, and coherence requires restraint. On a live buyer call, the architecture is invisible. What matters is whether the AI remembers what was said, responds consistently, and avoids contradicting itself as the conversation evolves. Reliability is therefore shaped as much by what a team leaves out as what it builds in.
Read what the room is reading.
New pieces for growth leaders, delivered as they publish. Unsubscribe whenever.








