Multi-agent AI: agents that work together, and why it goes wrong at the handoff
Several AI agents passing work to each other: when it pays off, why research says it fails so often, what A2A and MCP do and do not cover, and what a good handoff looks like in PHP.
Ask a group of AI agents to build Flappy Bird together. One agent makes the background, the second makes the bird, a third puts it together. There is a good chance you get a background in the style of Super Mario, a bird that does not fit it, and a third agent that cannot make anything of the pieces. The example comes from Cognition, the company behind the coding agent Devin, in a June 2025 post titled Don't Build Multi-Agents.
A little over a year later the industry is moving the other way. Google, Microsoft and AWS support A2A, an open protocol that lets agents from different vendors pass work to each other. Anthropic has a lead agent in its own research feature split the work across subagents that search in parallel, and measures a large gain from it.
Both are right. Collaborating agents work well for one kind of work and badly for another, and the difference almost always lies in the moment one agent hands something to another. Sending that message has been standardised. What goes into it, and who checks what comes back, is still your job.
The facts and figures in this article were checked on 29 September 2026. Research results, company claims and my own assessment are marked separately throughout. The code examples are in PHP.
What is multi-agent AI?
In multi-agent AI, several AI agents work on one assignment, each with its own task, its own instructions and often its own tools. One agent splits the work or hands part of it off, another carries it out and returns a result. An agent in this sense is a language model that picks its own next step in a loop: call a tool, read the result, continue.
That one term covers two situations that have little to do with each other in practice:
| Within one system | Between systems | |
|---|---|---|
| Who builds the agents | you, or a single vendor | different teams or companies |
| Example | a lead agent with subagents, as in Claude Code or Anthropic's Research feature | your application has another party's agent carry out a task |
| What they know of each other | a lot: same framework, often shared context | only what is in the message |
| Standard needed? | no | yes: A2A (Agent2Agent) |
| Typical failure | duplicated work, conflicting choices | missing context, trust, who may do what |
The term agent interoperability refers to the right-hand column. Most problems in this article apply to both.
When collaborating agents pay off
The clearest public figure comes from Anthropic. In June 2025 the company described how its Research feature works: a lead agent (Claude Opus 4) breaks a question down and starts subagents (Claude Sonnet 4) that each investigate a part in their own context window. On Anthropic's own internal research evaluation, that setup scored 90.2 percent better than Claude Opus 4 working alone.
The example Anthropic gives shows what kind of work this is: find all board members of the IT companies in the S&P 500. One agent working through them company by company gets stuck. Ten agents that each take a batch of companies are done before the first is halfway. According to Anthropic, searching in parallel made research on complex questions up to 90 percent faster.
There is a bill attached. According to the same publication, an agent uses about four times as many tokens as a normal chat, and a multi-agent system about fifteen times as many. On the BrowseComp benchmark, the number of tokens used explained 80 percent of the differences in performance on its own. Part of the gain, in other words, is simply more compute.
That is the first question for any multi-agent plan: does the work split into pieces that do not depend on each other, and is the result worth fifteen times the tokens?
Anthropic also says where it does not work: tasks in which every agent needs the same context, or in which the agents depend on each other a lot. According to the company, most coding tasks belong there, because code contains fewer truly parallel pieces than research. That is Cognition's Flappy Bird problem, seen from the other side. Cognition put it as two rules: share the full context rather than individual messages, and remember that every action carries a decision that can clash with another agent's.
Why multi-agent systems fail
Researchers at UC Berkeley analysed 1,642 runs of seven open-source multi-agent frameworks, including ChatDev, MetaGPT, Magentic-One and OpenManus. Those systems failed in 41 to 86.7 percent of cases. The paper, Why Do Multi-Agent LLM Systems Fail?, appeared at NeurIPS 2025 and sorts the failures into fourteen modes in three groups.
The most common failures are not exotic. An agent repeats steps that were already done (15.7 percent). An agent reasons towards A and then does B (13.2 percent). The system does not notice the task is finished and carries on (12.4 percent), or does not stick to the assignment (11.8 percent). In 6.8 percent an agent continues on a wrong assumption when it could have asked a question.
Two findings are the most useful for a developer.
Verifying is not the same as adding a reviewer agent. Many built-in verifier agents only checked whether the code compiled or whether TODO comments were left, even when they were told to check thoroughly. A chess program from ChatDev passed its review that way and then turned out not to follow the rules of chess. When the researchers added a check against the actual goal of the task, 15.6 percent more tasks succeeded. Clearer role descriptions alone gave 9.4 percent.
A protocol does not solve misalignment. The researchers name MCP and A2A: those standardise the format of messages between tools and agents from different makers. But the misalignment failures they saw also occurred between agents in the same framework talking to each other in plain natural language. The problem is that an agent misjudges what the other agent needs to know. A tidy message format does not change that.
A caveat: the measurements were made with models from 2024 and 2025, such as GPT-4o and Claude 3.7 Sonnet. Newer models make fewer of these mistakes. My assessment is that the classification itself stays useful, because most failures originate in how the system is designed, not in the model.
A2A and MCP: what the protocols cover
There are two protocols worth knowing, and they do different things.
MCP (Model Context Protocol) covers how an agent uses tools and data: a database, an API, a folder of files. Anthropic introduced it in November 2024. Since 9 December 2025 it sits under the Linux Foundation's Agentic AI Foundation, together with Block's goose and OpenAI's AGENTS.md; that foundation had 247 members on 13 August 2026. For PHP there is an official SDK, mcp/sdk, built by the Symfony team together with the PHP Foundation. How to connect MCP safely to your own systems is covered in my article on MCP.
A2A (Agent2Agent) covers how agents find each other and hand work to each other, including when they run at different companies. Google presented it in April 2025 and transferred it to the Linux Foundation on 23 June 2025. IBM merged its own Agent Communication Protocol into A2A in August 2025. Version 1.0 was released on 12 March 2026. According to the Linux Foundation, more than 150 organisations supported the protocol in April 2026, and it is built into Azure AI Foundry, Copilot Studio and Amazon Bedrock AgentCore.
The A2A specification describes agents as opaque: one agent cannot see which model the other uses, which instructions or which tools. It sees a description of what the other can do, sends a message and gets a task back with a status. Between companies that is exactly right. It also means the receiver knows nothing that is not in the message, and that is the most important design rule for everything that follows.
An A2A handoff in PHP
A2A 1.0 has three transports: JSON-RPC, gRPC and plain HTTP with JSON. For a PHP application, HTTP with JSON is the simplest; you need nothing more than curl. PHP libraries for A2A exist, but they come from the community rather than from the A2A project itself. For this example I stick to the bare calls, so you can see what goes over the wire.
The example is made up to show the technique: a PHP web shop has product texts translated into German by a translation agency's agent. The agency publishes an Agent Card at a fixed location, /.well-known/agent-card.json, with the agent's name, endpoints, authentication and skills:
{
"name": "Translation agent",
"description": "Translates product copy NL-DE using the client's terminology.",
"supportedInterfaces": [
{"url": "https://agent.translation-agency.example/a2a", "protocolBinding": "HTTP+JSON", "protocolVersion": "1.0"}
],
"capabilities": {"streaming": false, "pushNotifications": true},
"skills": [
{"id": "product-nl-de", "name": "Product copy NL-DE", "description": "Title and description, HTML stays intact.", "tags": ["translation", "e-commerce"]}
]
} The web shop reads that card, picks the endpoint and sends a message. A message consists of parts: text, a file or structured data. Here the instructions go as text and the product as JSON:
<?php
function a2aPost(string $baseUrl, string $path, array $body, string $token): array
{
$ch = curl_init(rtrim($baseUrl, '/') . $path);
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_TIMEOUT => 30,
CURLOPT_HTTPHEADER => [
'Content-Type: application/a2a+json',
'A2A-Version: 1.0',
'Authorization: Bearer ' . $token,
],
CURLOPT_POSTFIELDS => json_encode($body, JSON_THROW_ON_ERROR),
]);
$raw = curl_exec($ch);
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
if ($raw === false || $status >= 400) {
throw new RuntimeException("A2A call failed (HTTP $status)");
}
return json_decode($raw, true, flags: JSON_THROW_ON_ERROR);
}
// 1. Find the agent: the Agent Card lives at a fixed location
$card = json_decode(
file_get_contents('https://agent.translation-agency.example/.well-known/agent-card.json'),
true, flags: JSON_THROW_ON_ERROR
);
$endpoint = null;
foreach ($card['supportedInterfaces'] as $interface) {
if ($interface['protocolBinding'] === 'HTTP+JSON' && $interface['protocolVersion'] === '1.0') {
$endpoint = $interface['url'];
break;
}
}
// 2. The handoff: instructions as text, the product as data
$response = a2aPost($endpoint, '/message:send', [
'message' => [
'messageId' => bin2hex(random_bytes(16)),
'role' => 'ROLE_USER',
'parts' => [
['text' => $instructions],
['data' => $product, 'mediaType' => 'application/json'],
],
],
], getenv('TRANSLATION_AGENCY_TOKEN')); What comes back is a task with an id and a status. Sometimes an agent replies with a plain message instead of a task; the specification allows that too. A task moves through a fixed set of states, and your code needs to handle all of them, not just the one where everything went well:
<?php
$task = $response['task'] ?? null;
if ($task === null) {
// The agent answered with a plain message instead of a task
return handleDirectMessage($response['message']);
}
$result = match ($task['status']['state']) {
'TASK_STATE_COMPLETED' => verifyTranslation($task['artifacts'] ?? [], $product),
'TASK_STATE_INPUT_REQUIRED' => askEditor($task), // the agent asks a question
'TASK_STATE_AUTH_REQUIRED' => throw new RuntimeException('Agent requires extra authorisation'),
'TASK_STATE_SUBMITTED',
'TASK_STATE_WORKING' => schedulePoll($task['id']), // fetch later with GET /tasks/{id}
default => logFailure($task), // FAILED, CANCELED or REJECTED
}; INPUT_REQUIRED is the most interesting state. It is the place in the protocol where an agent can say: I don't know this, tell me. In the Berkeley study, 6.8 percent of failures were agents that guessed instead of asking. The protocol makes asking possible. Whether the agent actually asks depends on the instructions you send along.
The handoff as a contract
Anthropic describes what a lead agent has to give a subagent: an objective, an output format, guidance on which tools and sources to use, and clear boundaries for the task. Without those four, their early versions went wrong: agents duplicated each other's work, left gaps, or started fifty subagents for a simple question. The same four points work for a handoff between companies:
<?php
$instructions = <<<TXT
Objective: translate title and description into German for a German web shop.
Use the informal du form and do not add sales copy.
Output: JSON with exactly the fields "title" and "description".
The same HTML tags as the input, in the same order.
Sources: use the terminology list in the data. Do not translate brand
names or product codes.
Boundaries: do not shorten, do not add. If you are unsure about a term,
ask a question instead of guessing.
TXT; The last sentence ties the instructions to INPUT_REQUIRED: you give the agent permission to ask, and your code knows what to do with the question.
The second half of the contract is on your side: check what comes back in plain code, on things you can measure. Not with a second agent that says it looks fine, because then you have only moved the verification problem from the study.
<?php
function verifyTranslation(array $artifacts, array $product): array
{
$data = $artifacts[0]['parts'][0]['data'] ?? null;
$keys = is_array($data) ? array_keys($data) : [];
sort($keys);
if ($keys !== ['description', 'title']) {
throw new UnexpectedValueException('Translation is not in the agreed format');
}
// Same number of HTML tags as the original
if (substr_count($data['description'], '<') !== substr_count($product['description'], '<')) {
throw new UnexpectedValueException('HTML structure differs from the original');
}
// Product codes must come back verbatim
foreach ($product['codes'] as $code) {
if (!str_contains($data['title'] . ' ' . $data['description'], $code)) {
throw new UnexpectedValueException("Product code $code is missing from the translation");
}
}
return $data;
} This function does not establish whether the German is good. It does establish that the format is right, the markup is intact and no product code went missing. Those are exactly the errors a person easily overlooks when proofreading, and that code never overlooks.
What A2A and MCP do not cover
The protocol delivers the message. Two things remain your responsibility.
Security
On 9 September 2026, researchers at Purdue University published A2ABreak, a systematic analysis of the A2A specification. They found eleven new vulnerabilities, including injecting context across different clients through unprotected identifiers, harvesting credentials because identity gets lost along a chain of delegated tasks, and leaking data through a rogue agent claiming capabilities that nobody verified.
A2A version 1.0 includes signed Agent Cards, which let you check that a card really comes from the domain it claims; the specification says clients should do so. My advice goes one step further: treat everything another agent returns the way you treat user input. An artifact can contain instructions that your own agent then carries out. What an agent needs around it in terms of security and oversight is covered in the article on agent harnesses.
Who decides
In June 2026, Richard Kang and Yudho Diponegoro compared five protocols for collaborating agents, including MCP, A2A and ACP, on six aspects of decision-making: membership, deliberation, voting, preserving dissent, escalating to a human, and being able to replay afterwards what happened. No protocol supported all of them, and voting and preserving dissent were missing everywhere. Their conclusion: this is not a missing feature within the protocols but a layer above them that does not exist yet.
For a PHP application that means: who may approve what an agent proposes is something you decide in your own code. Log the task id, what you sent and what came back for every handoff, so you can trace afterwards which agent made which decision.
Where to start
Start with one agent. Split only when the work really falls apart into pieces that do not depend on each other, such as searching many sources at once, and budget for a multiple of the tokens. Only use A2A when the other agent runs outside your own system; within one application, a function call or a queued job is simpler and easier to test.
And treat every handoff as a contract with two sides: instructions with an objective, output, sources and boundaries, and a check in your own code on what comes back. The protocol delivers the parcel. What is inside remains your job.
Sources
- Anthropic, How we built our multi-agent research system (13 June 2025)
- Walden Yan (Cognition), Don't Build Multi-Agents (12 June 2025)
- Cemri et al., Why Do Multi-Agent LLM Systems Fail? (NeurIPS 2025)
- A2A project, Agent2Agent Protocol Specification (version 1.0)
- Linux Foundation, A2A Protocol Surpasses 150 Organizations (9 April 2026)
- Linux Foundation, Formation of the Agentic AI Foundation (9 December 2025) and 57 new members (13 August 2026)
- MCP blog, Announcing the Official PHP SDK for MCP (5 September 2025)
- Lotfi et al., A2ABreak: Systematic Security Analysis of the A2A Protocol (9 September 2026)
- Kang and Diponegoro, Governance Gaps in Agent Interoperability Protocols (30 June 2026)