The success-failure ratio of Multi-agent LLM systems often leans more to the right side despite all the enthusiasm about the technology. Researchers at UC Berkeley analysed Multi-Agent System Failure Taxonomy (MAST) data gained from 1600-plus execution traces across seven production-grade agent frameworks. Analysis disclosed a 41% to 86.7% failure rate, revealing a vast gap irrespective of the AI model used for agents. This diversity proves that there is no universal solution for all the failure modes of MASs and shows the necessity of analysing the problems that hinder the effectiveness of multi-agent systems.
There are three categories of structural causes behind MAS breakdown: failure modes in system design, failure modes in task verification, and misalignment between agents.
However, the taxonomy presents a grounded theory analysis of execution traces across seven MAS frameworks, enabling engineering teams to track where multi-agent systems actually break.
This article will talk about why do multi-agent LLM systems fail, 14 modes, the production data behind each one, and the architectural patterns that prevent them.
Table of Contents
ToggleA multi-agent system is an architecture where multiple LLM-powered agents coordinate to complete a task that no single agent could reliably finish alone. Each agent acts with a distinct role, tool set, or scope of authority, where one might plan, another might execute code, and a third might verify output. The layers and structural complexity of multi-agent systems are where things break down.
Coordination is the main source of power for MAS, and it is also the reason for its fragility. During each handoff between agents, loss of context and reinterpretation of the instruction can happen, compounding errors. Adding more agents may increase the potential for high performance, but it also contributes to the number of failure points in the system.
Pro tip: System design issues, inter-agent misalignment, and verification gaps actually originate at agent-to-agent handoff. So treat them as a potential failure boundary, not a formality.
The MAST study (Multi-Agent System Failure Taxonomy) is the first empirically grounded classification of why multi-agent LLM systems fail, built from actual execution traces rather than intuition. Researchers collected 1,642 execution traces across seven widely used multi-agent frameworks, including MetaGPT, ChatDev, AppWorld, Magentic-One, and AG2.
The comprehensive MAST-Data reveals several critical insights about failure patterns in multi-agent systems. Six expert annotators independently labeled each of the 150 traces against the taxonomy, reaching a Cohen’s Kappa agreement score of 0.88.
Every one of the seven frameworks tested failed between 41% and 86.7% of the time, regardless of the underlying model. Many MAS frameworks exhibit alarmingly low performance gains, often minimal compared to simpler single-agent systems.
The 14 failure modes cluster into three categories:
| Category | Share of Failures | What It Covers |
|---|---|---|
| System Design Issues | 41.8% | Task misinterpretation, role confusion, step repetition, lost context, unrecognized task completion |
| Inter-Agent Misalignment | 36.9% | Context collapse, error compounding, format mismatches, conflicting objectives |
| Task Verification Failures | 21.3% | Shallow or incorrect verification, premature termination |
Identify where your agents break, close the gaps at every handoff, and ship multi-agent workflows that hold up in production.
This issue drives more multi-agent systems failures than any other category, accounting for 41.8% of all failures. The MAST researchers observed that system design failure is spread across five distinct modes, and each one is a structural gap in how the system was specified before execution ever started.
Real example from research: Researchers asked ChatDev to build a Wordle game that selected a new random five-letter word daily instead of pulling from a fixed dictionary. Even with a more explicit prompt, ChatDev still produced code with a fixed list and new errors. The system never adapted, even after the constraint was spelled out twice.
Why it happens: Agents interpret task specifications locally, the way a single LLM model interprets a prompt, without cross-checking against the full intent. Agents drift toward the most common pattern in their training rather than the one actually requested which becomes the reason for why do multi-agent LLM systems fail.
How to prevent it? Ensure constraint checklists are validated against output, not just restated in the prompt, and the orchestrating agent has to reinterpret at every step. Encode task constraints as structured, machine-checkable specifications.
Real example from research: In ChatDev, the CPO agent could terminate a conversation without securing the CEO agent’s consent, a direct violation of the designed reporting hierarchy in the system.
Why it happens: Role boundaries in most multi-agent systems live in prompt text, not enforced permissions. Any agent can act outside its defined scope because nothing in the architecture actually stops it.
How to prevent it? Enforcing that only the CEO agent could finalize a conversation increased task success by 9.4% in ChatDev, with no change to the underlying model.
Real example from research: Planner agent in a HyperAgent trace, working on a matplotlib bug fix, repeating the identical diagnostic reasoning twice.
Why it happens: It typically stems from agents losing track of what they’ve already tried, especially in frameworks without a persistent, checkable execution log.
How to prevent it? Maintain a shared, agent-readable record of completed steps that each agent checks before acting.
Real example from research: MAST researchers define this mode as unexpected context truncation when an agent disregards recent interaction history and reverts to an earlier conversational state.
Why it happens: Long-running multi-agent workflows accumulate conversation history faster than most context windows can hold, and truncation happens without the agent noticing it.
How to prevent it? Externalize state to a persistent memory store rather than relying on the raw conversation window, and summarize checkpoints an agent can reload.
Real example from research: In an AG2 math session, the assistant correctly identified that a problem lacked enough information to solve and stated this three separate times. But the orchestrating agent still responded, “Continue. Please keep solving the problem until you need to query,” every time, unable to recognize that the task had already reached a valid terminal state.
Why it happens: Unawareness of termination specifications usually happens because success and failure criteria were never defined as explicit, checkable conditions.
How to prevent it? Do not leave the termination conditions on the agent’s judgment about when it feels finished; rather, define them explicitly and in testable criteria.
The second-largest category in the taxonomy for multi-agent systems failures is misalignment between agents, contributing to 36.9% of all documented MAS errors. MAST researchers find this structural failure the hardest to fix, and standardizing message formats does not resolve the underlying problem:
Real example from research: A MAST-annotated trace shows that, in a login task, the required username field expects a phone number format, not a generic username. But both the phone agent and supervisor agent never communicate these critical details with each other, producing repeated failed logins.
Why it happens: Neither of the agents models what information the other one actually needs, relying on mismatched assumptions.
How to prevent it? Require agents to explicitly state their information dependencies before a handoff, not just their outputs. A receiving agent that has to confirm, “Here’s what I still need”.
Real example from research: In a chain of 5 agents, each having 95% reliability, the ultimate success rate is just around 77% and not 95%. It happens because reliability does not compound on an average basis in multi-agent systems but multiplicatively across the chain. The longer the chain, the higher the failure rates.
Why it happens: When an agent’s stated logic and its actual behavior diverge, a subtly incorrect output from a previous agent is treated as ground truth by the next agent in line.
How to prevent it? Place an adversarial verification agent after each primary agent, explicitly tasked with catching errors rather than passively reviewing.
Pro tip: Treat every agent handoff as a place errors can enter the system, not just a place work gets passed along.
Real example from research: In an analysis of the AgentVerse framework running on a Qwen-2.5-14B-Instruct backend, researchers found format errors. They did not get wrong answers, but the responses mismatched the required output structure.
Why it happens: Agents typically communicate through unstructured natural language rather than validated schemas. For instance, the planner produces YAML when the executor expects JSON and has no built-in mechanism to catch the mismatch before it cascades into a downstream failure.
How to prevent it? Validate every inter-agent message against a defined schema at the boundary, rejecting and retrying malformed handoffs rather than letting a receiving agent guess at intent.
Real example from research: Researchers modeling a retail deployment describe two agents optimizing against the same inventory. One is maximizing product fill rates, thus increasing reorder volume. While the other, minimizing cash tied up in stock, starts delaying purchase orders in response. Here, both have conflicting objectives and neither agent malfunctions individually, but the resource contention itself produces an escalating, unproductive loop.
Why it happens: When two agents believe they independently own the same resource, each acts rationally from its own objective while working directly against the other.
How to prevent it? Assign exactly one agent as owner of each resource or data store, and route any cross-agent dependency through an explicit shared-state layer with documented read/write semantics, never through two agents independently modifying the same target.
Task verification failures make up the smallest of the three MAST categories, split across three major error points:
Real example from research: A ChatDev-generated chess program passed every check its verification agent ran; however, the program still contained runtime bugs that broke actual gameplay because the verifier never checked the output against real chess rules.
Why it happens: MAST researchers found that many existing verifiers, despite being explicitly prompted to verify thoroughly, perform only superficial checks confirming code compiles or a response is well-formed.
How to prevent it? Define verification criteria tied to the actual task specification, game rules, business logic, functional requirements, not just execution success.
Real example from research: In a physics problem-solving trace, an agent produced an answer of 300 N/C that should have been mathematically rejected outright. Instead of catching the inconsistency, the Solver agent introduced an additional persona, a “senior physicist”, into the conversation, and the incorrect answer was never actually resolved or corrected.
Why it happens: A verifier can be present in the architecture and still lack the reasoning depth to catch a specific error class.
How to prevent it? Give verification agents a narrow, well-defined check to perform bounds checking, unit validation, rule compliance, rather than an open-ended instruction to “review the answer.”
Real example from research: AppWorld showcased a disproportionate concentration of premature terminations relative to the other systems studied. Researchers attribute this pattern to its star-topology architecture, where no predefined workflow makes it clear to any agent when the task has actually reached a valid end state.
Why it happens: Without an explicit, externally checkable definition of “done,” an orchestrating agent has to infer completion from conversational cues.
How to prevent it? Define completion as a structural property of the workflow, including a required set of outputs or a fixed sequence of steps.
Most multi-agent systems fail under the veil, and only manual inspection can surface them before reaching downstream. Unnoticed errors from automated monitoring not only cause broken outputs with low reliability rates, but also tamper with the budget line. 41%-87% failure rate indicates invisible errors that no system detects unless it becomes expensive in the following three compounding ways:
Agentic workflows consume roughly 1,000 times more tokens per task than a standard chat interaction. Each phase and action requires re-sharing of accumulated context under a multi-agent workflow, increasing the interaction cost roughly 30 times compared to a single linear exchange.
Debugging a multi-agent failure is a significantly lengthier process than finding an error in a single-agent process, as most defects happen at the time of handoff between agents. A broken output arrives as a single bad result with no visible clarity about which agent introduced the error and which one just passed it further in the chain. Without structured tracing, engineers are forced to manually reconstruct the trace step by step.
Shifting from a single-agent system to a multi-agent one doesn’t just double the budget; it multiplies it by 5 to 10 times. Developing and operating a multi-agent system demands expansive orchestration logic, failure handling, shared memory, and evaluation infrastructure.
None of these costs show up in a proof-of-concept demo, which is precisely the problem: a system that fails silently in a controlled pilot fails silently and expensively in production, and the bill arrives after the architecture is already load-bearing.
Traditional debugging assumes two things: that multi-agent systems follow a deterministic execution path and provide a clear point of failure. But the multi-agent LLM system breaks both of these assumptions by demonstrating that every run of the MAS workflow can follow a different path even with the same inputs, and failure in the final output often originates from several agents and several steps earlier.
This gap makes conventional logging inefficient, as the old methods depend on the discrete events that happen in isolated agent workflows. But a multi-agent LLM system distributes decision-making across autonomous components that coordinate without centralized control, so a single flat log gives fragments of five separate conversations with no causal thread connecting them.
Debug multi-agent LLM system with a conventional approach, and these three specific properties will make it worse:
The system executes end-to-end and returns a credible-looking answer that’s simply wrong. This is what most multi-agent breakdown looks like: not crashed output, but silent failure.
Root-caused failures in multi-agent systems frequently arise from inter-agent misalignment instead of a defect inside any single agent. Here, the real culprit is not the function, but the handoff, and traditional debugging mechanisms are not built to trace multi-agent failure.
An error introduced at step 3 of a workflow can remain invisible until it reaches the final output at step 20. Till then, the original failure may get buried under dozens of handoffs and intervening agents.
To fix these debugging issues, multi-agent observability is required to trace failure across the full execution path, not just at the individual agent level. Structured trace trees, semantic context, and cross-agent correlation are the three pillars on which effective visibility relies in multi-agent systems.
Along with knowing why do multi-agent LLM systems fail, it is also important to understand how to build systems that don’t fail. Every multi-agent system failure traces back to the same root cause: framing a team by wiring up several agents, and then discovering any error through the broken output. An architecture-first approach inverts that order and configures the failure before writing any single agent prompt. Only after mapping the potential failures, the following five layers are built in sequence:
Before selecting any framework, map your planned workflow against the 14 MAST failure modes and find which ones your specific task is most exposed to. Framework choice greatly impacts downstream failure exposure, but not vice versa. Thus, optimize LangGraph, CrewAI, AutoGen, or any other framework first before making a final pick.
Three orchestration patterns dominate production systems, and each has a distinct failure signature:
| Pattern | Structure | Strength | Risk |
|---|---|---|---|
| Sequential Pipeline | Agent A → Agent B → Agent C | Predictable, easy to trace | Errors compound linearly; one bad handoff breaks everything downstream |
| Hierarchical (Supervisor) | Supervisor → {Agent A, Agent B, Agent C} | Centralized verification point, clear ownership | Supervisor becomes a bottleneck and a single point of failure |
| Star / Decentralized | Agent A ↔ Agent B ↔ Agent C (peer-to-peer) | Flexible, no single bottleneck | Highest exposure to premature termination – AppWorld’s star topology is the framework MAST found most prone to this failure |
A hierarchical pattern with a supervisor agent responsible for both routing and final verification directly closes two of the highest-frequency failure modes: role confusion and shallow verification using construction rather than convention.
Every inter-agent message should validate against a defined schema before the receiving agent acts on it, not after something breaks downstream:
{
"from_agent": "planner",
"to_agent": "executor",
"task_id": "task_0142",
"payload": {
"action": "fetch_records",
"params": { "table": "orders", "filter": "status=pending" }
},
"context_summary": "Retrying after executor reported empty result set",
"schema_version": "1.2"
}
Developing an observability layer first means every agent action is captured as a span in a structured trace tree from the very first test run rather than reconstructing failures from incomplete logs after the fact:
trace_id: run_88231
├── span: planner.decide_next_step [420ms]
├── span: executor.call_tool(fetch_records) [1,240ms]
│ └── error: schema_validation_failed
├── span: executor.retry(fetch_records) [890ms]
└── span: verifier.check_output [310ms]
└── result: incomplete — missing 3 required fields
End point verification invites the shallow-check problems most multi-agent system failures arise from, confirming the delivery of output without analysing its correctness. Treat verification instead as a dedicated layer with narrow, task-specific checks positioned at every agent boundary, not just at the end of the pipeline.
Architecture-first design, structured tracing, and verification at every handoff.
The NineHertz is an AI-native engineering firm built around a Build, Run, and Evolve framework that treats reliability as a first-class design constraint rather than a post-launch fix. Most teams learn the architecture work the expensive way, by shipping a pilot that works in the demo and fails in week three of production. The NineHertz’s framework maps directly onto the architecture-first approach:
Common multi-agent LLM system failures occur due to architecture problems, not model quality. Three structural reasons account for most of the MAS errors, including system design issues, misalignment between agents, and verification layers.
Researchers analyzing 1,642 execution traces across seven production-grade multi-agent frameworks found failure rates between 41% and 86.7%, depending on the framework, with the pattern holding across different underlying models including GPT-4, Claude 3, and Qwen2.5.
Error compounding happens when reliability multiplies rather than averages across a chain of agents. Five agents each running at 95% individual reliability produce roughly 77% end-to-end success, not 95%, because every downstream agent inherits whatever the upstream agent got wrong.
Costs compound across three areas: compute, per-interaction cost, and build cost. For instance, multi-agent architectures typically run 5 to 10 times more expensive to build and operate than single-agent systems once orchestration, failure handling, and evaluation infrastructure are included.
MAST (Multi-Agent System Failure Taxonomy) is an empirically grounded classification of 14 distinct multi-agent LLM failure modes, developed by researchers at UC Berkeley from grounded-theory analysis of 1,642 annotated execution traces across seven frameworks.
No. MAST researchers tested the same multi-agent frameworks across multiple underlying models, including GPT-4, Claude 3, Qwen2.5, and CodeLlama. They found failure rates stayed high across all of them, and thus concluded that how the system is architected decides the failure rate and not the quality of the model powering any one agent in it.
Kapil Kumar co-founded The NineHertz and has spent over a decade building teams, products, and businesses across global markets, evolving from writing code and delivering projects to architecting systems that scale under real-world pressure. As Co-Founder and Chief Growth Officer, his expertise centers on AI consulting, product strategy and planning, and go-to-market strategy, paired with strong technology leadership and a proven ability to build and scale technology teams.
Kapil’s approach is defined by execution-focused leadership that transforms strategy into measurable business outcomes through clarity, timing, and disciplined delivery. He combines deep technical expertise in web and mobile application development with a business-first lens, helping organizations use technology as a practical lever for efficiency, control, and long-term growth. His leadership has been instrumental in shaping The NineHertz into a resilient, quality-driven organization built to scale alongside its clients.
Key Takeaways Nearly 40% of respondents in India report significant or full AI use, compared with 28% globally. Hourly rate…
Key Takeaways Traditional SCRMs are ineffective for modern supply chain enterprises, because they fail to detect risks in time, offer…
Key Takeaways AI supply chain risk monitoring uses real-time data and predictive models to flag disruptions before they escalate into…
Take a Step forward to Turn Your Idea into Profit Making App