Intermediate Any tool News

Agent security: attackers are running agents too

OWASP's incident data puts agents where the damage lands, and Anthropic's threat report shows attackers using agents of their own.

The incident figures are OWASP's own, and the case studies are Anthropic's.

Two things arrived in the evidence this summer. OWASP, the security community whose top-ten lists much of the industry works from, checked its AI risk list against real incidents for the first time, and the incidents point at agents. And Anthropic’s September threat report describes attackers running agents of their own, doing work that a year ago took a team.

Where the damage is landing

OWASP re-issued its Top 10 for LLM applications on 4 August 2026. The list still rests on a vote of practitioners, but this year a quarter of the weight came from 6,639 real incidents. The biggest move was excessive agency, meaning an AI given more permissions or more freedom to act than its task needs, which climbed to third place. OWASP’s reason: “the vote and the record agree that agentic deployments are where the damage is landing”.

Prompt injection, where instructions hidden in something the AI reads get obeyed as if you had typed them, kept first place. The widest gap was on misinformation: voters put it near the bottom and the incident record put it near the top. OWASP’s explanation is worth reading twice: “When a model’s fluent, confident output drives a decision or a tool call, a wrong answer turns into a wrong action.”

Attackers have agents too

Anthropic’s report covers misuse of Claude it detected and disrupted between December 2025 and August 2026. Its headline is that AI “has collapsed the labor and tooling gap” that used to separate state-backed hacking teams from individuals. Most of the cyber cases used AI to carry out or co-ordinate the attack itself, with several agents working together to find a way in and take data out. People still chose the targets and reviewed what was taken.

The part that lands on ordinary organisations is theft of AI access itself. An API key is the password that lets software use an AI service, and the report found stolen keys and login sessions being sold through brokers and fraudulent reseller networks, then used to run further attacks. Its advice is for organisations to treat AI keys “with the same level of seriousness as they do production credentials”, meaning the keys that run live systems, and to buy AI access “only through authorized channels”, never through a discount that routes your traffic and credentials via an unknown middleman.

Treat every AI API key as a production password. Anthropic’s September report found stolen keys being resold and used until their usage ran out.

The strongest cyber skills are being rationed, mostly

Anthropic’s newest top model comes in two versions. The standard Claude Fable 5.1 can now be used to find software vulnerabilities, “though not to develop exploits for them”. Claude Mythos 5.1 is the same model with “more permissive safeguards” for vetted cyber defenders and life scientists, and it is “only available to a set of US organizations”. The gate has a gap. Z.ai published the weights of GLM-5.3 for anyone to download, and says of its own model that “cyber capability developed faster than we expected”. It claims the model more than doubles its predecessor on benchmarks for exploiting software flaws. Nobody can put a gate in front of weights that are already public.

A standard for the brakes

OWASP also published version 0.1 of an Agent Control Standard in September. It lets a separate “Guardian” agent check what an AI agent is about to do and allow, block or change it first, with an audit trail. Read the default before adopting it: the Guardian fails open, so one that “crashes, hangs, or is unreachable stops governing, silently”. The project’s own reference Guardian also skips the message signing the standard requires, so treat it as a demonstration. Anyone deploying the standard at work should set it to block on failure.

If you run agents where you work, the useful checks this month are short. Know which agents exist and whose keys they hold. Then look for any agent that combines untrusted input, such as email or web pages, with sensitive data and the power to act or send things out. OWASP recommends Meta’s “Rule of Two” here: an agent with all three needs a person to approve each action. For the July picture, see the earlier round-up.

Sources

Everything above was checked against these on 26 Sept 2026. Providers change things without notice. If a detail matters to a decision, follow the link.

  1. OWASP GenAI LLM Top 10 2026: Letter from the Project LeadsOWASP GenAI Security Project · 4 Aug 2026
  2. LLM01:2026 Prompt InjectionOWASP GenAI Security Project · 4 Aug 2026
  3. Detecting and countering misuse of AI: September 2026Anthropic · 10 Sept 2026
  4. Introducing Claude Fable 5.1 and Claude Mythos 5.1Anthropic · 1 Sept 2026
  5. GLM-5Z.ai on GitHub
  6. Agent Control StandardOWASP GenAI Security Project

Search