When people worry about AI agents, they picture a model that goes rogue. Security researchers are finding a plainer problem. The weak spot is often the plumbing around the model. Call it the control plane: the gateways, tool connectors, settings, and permissions that decide what an agent can reach and do.
This matters because agents can act. If an attacker takes over that control layer, they do not need to fool the model. They can borrow the agent's access to code, files, and cloud accounts.
What Is Actually Being Attacked
Most of the attention is on MCP, the Model Context Protocol. Anthropic released it in November 2024 as a standard way for agents to connect to tools and data. Many AI coding tools and agent platforms now support it.
Researchers have found real flaws. In April 2026, OX Security reported that the official MCP software kits can run a command on the host computer before checking whether it belongs to a real MCP server. OX says an attacker could reach this through prompt injection, a changed settings file, or a poisoned tool listing. It showed working attacks on six live platforms and estimates up to 200,000 vulnerable instances. Those numbers are OX's own estimates.
OX also showed a zero-click attack on the Windsurf coding tool. If the tool opened a booby-trapped web page or project file, hidden text could change its settings and add a malicious server. That bug is tracked as CVE-2026-30615. Cursor and Claude Code needed some user action to be exploited.
Other flaws sit in gateways. Researchers at Horizon3.ai showed that a bug in LiteLLM, a popular open-source AI gateway, could be chained with a second bug to run code without logging in. Security alerts describe it as actively exploited. OX also found that nine of eleven public MCP marketplaces accepted a harmless test submission with no review.
Counts vary by tracker. One advisory counts more than 40 vulnerabilities in MCP software between January and April 2026. Another tally counts 14 by July. Treat the totals as rough.
What Is Proven and What Is Not
Proven: these flaws have official CVE numbers, which are public IDs for security bugs, and many have been patched. Researchers demonstrated the attacks on real systems, not just on paper.
Disputed: who is responsible. Anthropic told OX that the behavior is intentional. It updated its security guidance but did not change the design. Its position is that developers must limit which commands are allowed. OX argues that asking every developer to get this right will fail at scale. Neither side has settled it.
Unproven: the fixes being sold. Companies such as Palo Alto Networks and Astrix now sell "agent control planes." These tools give each agent short-lived credentials and a way to cut off access fast. That is a sensible idea, but I found no independent tests showing these products stop the attacks above.
Why It Is Hard to Defend
A researcher survey on arXiv says agent attacks are "compositional." An attack can pass through several parts of a system, so a defense in one place may never see it. Another research team argues agents should be secured like operating systems, with every action checked by a trusted layer. That is a research idea, not a finished product.
In plain terms, the model reads text it cannot fully trust, and that text can end up changing the settings that control it. Normal software keeps those two things apart. Many agent tools do not.
Security researcher Yotam Perkal described a common trap. When a company adds MCP to an existing app, the new endpoints get the app's full powers but not always its security checks.
What Could Happen Next
In the near term, rules and standards are arriving. OWASP published a Top 10 list of agent risks in December 2025. NIST launched an AI Agent Standards Initiative in February 2026 and released a concept paper on agent identity. A concept paper is a proposal, not a standard. On May 1, 2026, agencies including CISA and the NSA released joint guidance on securing agentic AI, according to a Cloud Security Alliance briefing.
The longer-term picture is speculation. Agent identity could become as routine as company logins. Or a large breach could force stricter rules. The open question is who sets safe defaults: the protocol makers, the tool builders, or the companies that deploy agents. Until that is settled, the control plane will stay an attractive target.