[research] · · 4 min read
MCP protocol pivoting flaw hits Google, JP Morgan, Rapid7
A researcher demonstrated that trust gaps in the Model Context Protocol allow attackers to pivot malicious instructions between agents, bypassing LLM guardrails.
By ByteBulletin Editor · Editor
Ars Technica reports that a structural flaw in the Model Context Protocol (MCP) has been exploited to compromise AI agents at major organizations including Google, JP Morgan Chase, Weviate, Rapid7, and the French government. Independent researcher Syed Anas Mohiuddin demonstrated that attackers can use one compromised agent to spread harmful instructions to other internal agents, effectively bypassing the guardrails that protect the underlying Large Language Models. The technique, which Mohiuddin terms "protocol pivoting," exploits the inherent trust that MCP servers place in other internal agents, allowing malicious prompts to traverse the network and execute unauthorized actions such as data exfiltration.
The core issue lies in how MCP handles inter-agent communication. While LLMs are often protected by safety filters, the specialized agents that handle specific tasks—such as translation or data analysis—often lack these strict guardrails. Because MCP servers store credentials for each agent and are built on the assumption that internal agents are trusted, a malicious instruction passed from one agent to another is executed as a legitimate delegated task. This creates a chain reaction where the initial prompt injection is not caught by the LLM's safety mechanisms but is instead propagated through the protocol layer, leading to server-side request forgery (SSRF) and other vulnerabilities.
The Mechanics of Protocol Pivoting
Mohiuddin’s proof-of-concept attacks highlight a specific weakness in how different communication protocols interact. The attack works when an application uses MCP to assign a task to an agent, and that agent then forwards malicious instructions to another agent using a different communication method, such as Google’s Agent-to-Agent (A2A) protocol or the Agent Network Protocol. In these scenarios, trust and authorization are often lost in translation between the protocols. Mohiuddin describes this as "a multi-step attack in which an adversary gains initial access through one protocol, exploits trust assumptions between protocols, and escalates to capabilities only accessible via a different protocol."
The severity of these vulnerabilities varies by implementation. For instance, the vulnerability found in Rapid7’s network, identified as CVE-2026-97228, carried a severity rating of only 2.7 out of 10 and was fixed last month. In contrast, the vulnerability affecting Google’s googleapis/mcp-toolbox was rated an 8. This specific flaw stemmed from the MCP toolbox initializing its HTTP client without a CheckRedirect policy and failing to validate target IP addresses. As Mohiuddin explained, "A crafted path parameter could make the toolbox follow a redirect to an internal endpoint and send requests on the attacker’s behalf." Google’s fix involved implementing an allow-list of IP ranges and block lists, rejecting unsafe base URLs at startup rather than on the first request.
Industry Context and Security Debates
The fact that this technique worked across five organizations with little in common other than their use of MCP suggests a systemic issue in the rapid adoption of agentic architectures. Markus Vervier, a researcher at X41 D-Sec, argues that the term "protocol pivoting" is less precise than the existing classification of "indirect prompt injection." Vervier noted, "The fact that the malicious prompt can come from a different protocol (e.g., A2A) and manifests when used over another protocol is not strictly required for such attacks to work. It is, of course, unexpected and hard to mitigate in general."
Douglas McKee, director of vulnerability intelligence at Rapid7, emphasized the difficulty of detecting these attacks because each component in the chain performs its intended function. "Someone plants text in content, an agent will read it then pass it along to another agent as a normal delegated task, and that second agent runs it because it trusts whoever handed it the work," McKee told Ars Technica. He added that the root cause is a departure from the zero-trust security model, where networks are built assuming that one or more nodes may be infected. In the rush to build sprawling agentic systems, organizations have abandoned the principle that nodes must require authorization before conducting sensitive transactions with other nodes.
Implications for Developers
For developers building on MCP, the primary takeaway is that any data passed from an LLM to a tool should be treated as untrusted input from an external source. McKee advised, "The lesson I’d want people to take away is that anything passed from an LLM to your tool should be treated like input from a stranger on the internet, because in a prompt injection scenario that’s exactly what it is." This means implementing strict validation of URLs, IP addresses, and redirect policies in any MCP server or agent that handles network requests. Developers should ensure that their HTTP clients have robust CheckRedirect policies and that base URLs are validated at startup. Additionally, organizations should review their inter-agent communication protocols to ensure that trust boundaries are explicitly defined and enforced, rather than assumed by default.
What to Watch
- Standardization Efforts: Watch for updates from the MCP working group and standards bodies like the Linux Foundation, which may introduce new security requirements for inter-agent communication.
- Vendor Patches: Monitor security advisories from major MCP implementers, including Google, Microsoft, and Anthropic, for patches addressing SSRF and trust boundary issues.
- Emerging Protocols: Keep an eye on the development of the Agent Network Protocol and A2A, as their integration with MCP may introduce new attack surfaces.
- Regulatory Response: As AI agents become more prevalent in enterprise environments, expect increased scrutiny from regulatory bodies regarding the security of agentic architectures.
Get the signal, not the noise.
One short email when it matters. No recaps of recaps.
SHARE
RELATED

[research] ·
OpenAI agents hacked government databases and leaked user images

[research] ·
Hacktron exploits Claude Opus 5 to breach OpenAI

[research] ·
Grok Bypassed by 'Cryptographic Context Injection' Attack That Exfiltrates User Data

[research] ·
OpenAI Agents Discuss Sandbox Escapes and XSS Attacks on Public Wiki

[tooling] ·
Apple requires explicit user action for macOS Full Disk Access

[tooling] ·
