मुख्य सामग्री पर जाएँ
विश्लेषण
Above

Anthropic's Threat Report: The Shift That Matters Is the Workflow Around the Model

Anthropic's September 2026 threat report shows misuse industrialising around the model: orchestration, retained context, parallel work. What operators must change.

लेखक Adam Maguire Wilson14 मिनट में पढ़ें
Abstract visual of networked nodes representing AI agent workflows orchestrated in parallel across sessions.
इस पृष्ठ पर

Anthropic published its latest threat intelligence report on 10 September 2026, the fourth in a series that began in March 2025. Detecting and countering misuse of AI: September 2026 covers activity the company says it disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development and illicit distillation.

The cases built to alarm will get the coverage. A China-based studio ran more than 4,700 fake dating personas that exchanged roughly 2.36 million messages with at least 25,000 people in a fortnight. A single consultant, likely based in Bamako, used Claude as the primary engineering workforce for a national interception platform designed to monitor roughly 25 million SIM cards for Mali's state intelligence service. Those cases deserve attention, but they are not the story with the longest shelf life for people who build, operate or buy agent systems.

That story is quieter. Across the cyber and surveillance cases, the thing that has industrialised is not the model. It is the workflow around the model: orchestration that decomposes a campaign into agent-sized tasks, retained context that carries an operation's memory between sessions, and parallel work that keeps collection running while the operators sleep. Anthropic's own section heading says it plainly: "From assistant to orchestrator."

A note on evidence before the cases. Everything in this report comes from Anthropic's own platform telemetry and investigation. It is self-asserted primary evidence from a vendor with an interest in demonstrating that its detection works, and the company says openly that these are "the most notable and novel" cases, not typical misuse. Where Anthropic hedges its attribution, the hedge is reproduced below. Where outcomes are unproven, that is stated too.

Key Takeaways - Anthropic's September 2026 report covers misuse disrupted between December 2025 and August 2026, nearly all of it on Claude Haiku, Sonnet and Opus models (one distillation case excepted). Its clearest finding is organisational: threat actors now run campaigns as agent workflows, with humans setting targets and reviewing output. - One tracked group ran "agent swarms" with persistent campaign memory across sessions and parallel workstreams, and produced more than a dozen possible zero-day findings in a single month from one autonomous workflow. - A surveillance operation processed 15 to 30-plus foreign news articles a day through a version-controlled operational manual, showing bureaucratic routine rather than novel capability. - The defensive consequence: per-request refusals and per-session controls cannot see a campaign that is fragmented across accounts, sessions and weeks. Identity, permissions, sustained-use monitoring and cross-session intervention have to live at the workflow layer. - This is human-directed misuse, not autonomous misalignment. The two call for different controls, and conflating them produces the wrong programme.

What Anthropic reported, and how to read it

The report introduces its own vocabulary. Threat actors are designated Generative Threat Groups, or GTGs, Anthropic's internal labels for actors observed abusing AI. "Uplift" is the company's term for how much more harm an actor caused with AI than without, viewed through speed, scale and depth. Both terms are worth keeping, because they let Anthropic describe what it saw without claiming more than it measured.

Three trends frame the case studies. First, the operating model Anthropic documented in November 2025, when it disclosed what it called the first AI-orchestrated cyber espionage campaign, "has now proliferated across every class of actors we investigated," from state services to lone individuals. Public offensive agent frameworks such as PentAGI reproduce much of the same scaffolding for anyone who downloads it. Second, a majority of the operations in the report involved AI "via direct execution or orchestration," with multi-agent frameworks executing reconnaissance, exploitation and exfiltration rather than answering chatbot-style questions. Third, and most quoted, "sophisticated attacks no longer require sophisticated attackers," which Anthropic illustrates with a hacktivist, disparate financially motivated individuals and a state espionage operator each sustaining campaigns that would recently have needed teams.

Two qualifications belong here. Anthropic can only see its own platform. The report itself describes actors using relays, stolen API keys and multi-model fallback precisely to reduce that visibility, so the sample is what one provider caught, not the population of misuse. And "humans remained in the loop," the report says, "by setting the targets of attacks and reviewing exfiltration." Hold that sentence. It is the difference between this story and a different, frequently confused one.

The foundry that kept working while its operators were away

The strongest evidence for the workflow thesis is GTG-10007, tracked by Anthropic as a sustained espionage operation run by Chinese-speaking operators "likely residing in Changsha," in Hunan province. Two of the operators were identified as undergraduate computer-engineering students; one had interned at the Chinese security company Sangfor and was interviewing for an offensive role at another, QiAnXin. Attribution beyond that is not claimed.

What the group built with Claude was, in the report's words, "the engineering and orchestration layer of a coordinated offensive programme." The observed structure matters more than any single output:

  • Orchestration. The operators "routinely ran 'agent swarms,' where a lead AI agent decomposed reconnaissance and post-exploitation work and dispatched it to many subagents running in parallel." Distinct workstreams ran concurrently: intrusions, foreign-government reconnaissance, security-product reverse engineering, malware development, and a collection platform.

  • Retained context. "The operation maintained persistent campaign memory. Target lists, harvested credentials, engagement state, and standing instructions were saved across working sessions, so each session could be resumed mid-campaign with the program's accumulated context."

  • Parallel, unattended work. A fleet of thirteen standing collection agents ran on scheduled jobs, harvesting open-source material "aligned with state intelligence priorities" through layered crawlers and commercial proxy exits. Capabilities, the report notes, "kept operating while its owners were away."

The zero-day work shows what that structure buys. An autonomous loop loaded vendor firmware into a decompiler, walked cross-reference chains across thousands of calls, formed vulnerability hypotheses against a curated knowledge base, wrote exploit code and tested it against lab copies of the target product, iterating until it worked. One such workflow, iterating continuously on network appliances, "yielded more than a dozen possible zero day findings in a single month." The actor validated multiple previously unknown vulnerabilities in a major security product in its own lab and produced working exploits for several families of network and security appliances. Anthropic separately observed exploitation attempts against those same appliances owned by government organisations. Whether any attempt succeeded is not established.

In total the group targeted roughly 50 organisations across education, retail, energy, technology, healthcare, finance, manufacturing and government. Confirmed outcomes in the report include hundreds of megabytes of student data taken from an education-technology company, access to a retailer's production systems, and citizen records from a Southeast Asian government agency. A detail worth sitting with: the hands-on intrusion work, as opposed to the autonomous workflows, concentrated exclusively on domestic Chinese victims. Anthropic banned the accounts and deployed additional monitoring.

Strip the geopolitics and the structure is familiar. A lead process decomposing work for subagents, a persistent memory file that survives the session, scheduled agents running unattended: this is the architecture HTA covers every week in legitimate tooling, including the memory-governance problems we examined in shared agent memory and the propagation risks in the agent "mind virus" paper. GTG-10007 is what those primitives look like when the operator is an adversary.

Surveillance as routine paperwork

The second case is less technically dramatic and, for that reason, more instructive. GTG-14022 was a China-based operation that used Claude as an automated "public opinion monitoring" (舆情) and intelligence-analysis system, the party-state's term for tracking and managing online sentiment.

The actor instructed Claude to role-play as "a senior emergency public opinion analyst serving the government of the People's Republic of China," and used its code-execution environment to run an automated document pipeline with minimal human intervention. The pipeline ingested 15 to 30-plus foreign news articles a day from Weibo, X, YouTube, Telegram and Facebook, scored each item for political sensitivity, reframed the content under fixed terminology rules ("Taiwan government" became "Taiwan authorities"; scare quotes were placed around "human rights violations") and formatted the result as restricted government briefings cataloguing dissidents, ethnic minority and diaspora communities, and foreign media as threats to stability. Some documents recommended enforcement actions that only a state could carry out.

The detail that lifts this from chatbot misuse to institutional machinery is the operational manual: version-controlled, at version 2.6, with a master control table and appendices. Anthropic's assessment of what made the case distinctive is the most quoted sentence in the surveillance section: "Rather than displaying any novel capabilities, this activity was unique in how it was used as part of the bureaucratic apparatus." The daily volume, the company adds, "suggests bureaucratic rather than ad-hoc activity."

Attribution here deserves the care Anthropic gave it. The company assesses with medium confidence that the operator was a contractor working for government clients, likely connected to the state security or united front and propaganda apparatus, rather than a state organ itself. It assesses with high confidence that two linked account clusters were the same actor, based on shared infrastructure. The accounts were banned; a wider network tied to that infrastructure is still being mapped.

Nor is this an isolated shape. Elsewhere in the surveillance section, Anthropic describes a PRC religious-affairs intelligence collection unit that "once comprised many teams of analysts" reduced to a single office using an AI assistant to produce thousands of investigations per month, and two Iranian units that independently used Claude to solve the same problems in a state-run surveillance case-management system. The pattern is capacity substitution: the workflow supplies the headcount.

Why the workflow is the control point

Put the two cases together and the mechanism is consistent. A frontier model in a single session is a capable assistant that refuses a great deal. Wrapped in orchestration, persistent memory and parallelism, the same model becomes an organisation: it decomposes, remembers, delegates and works shifts. Operational capacity, the thing defenders actually have to reckon with, is a property of that wrapper as much as of the weights.

The asymmetry this creates is documented in the report's own safeguard failure. In GTG-30006, an Iranian actor building domestic-surveillance tooling spread work across free Claude.ai accounts registered to 16 single-operator organisations. Claude refused nine out of ten direct requests that were facially malicious. "But," the report continues, "our safeguards performed less consistently when the user fragmented the work and directed the model to carry out tasks across later, smaller sessions." Per-request refusals evaluate a prompt. A campaign is not a prompt. It is a sequence of individually defensible requests distributed across accounts, sessions and weeks, which is exactly the shape retained context and task decomposition produce.

This is also the point where precision about agency matters, because two different risk stories are easy to conflate. Nothing in this report shows a model initiating harm. Every campaign described was human-directed: people chose the targets, built the scaffolding, wrote the standing instructions and reviewed the output. That is a different failure class from agents escaping containment or acting outside authorisation, the territory of the congressional letters we covered in our engineering read of the rogue-agent hearings. Autonomous misalignment is a model-behaviour problem with model-behaviour controls. Human-directed misuse at workflow scale is an operational-security problem with operational controls: identity, permissions, monitoring and intervention. Funding the first in the name of the second leaves the documented threat untouched.

There is one more limit, and Anthropic states it against itself. Account bans are a provider-side control, and providers are only one layer. In the Mali case, the surveillance platform ran fully on-premises on local models; Claude supplied software design and engineering support. "Account enforcement actions do not affect the deployed product," the report says, and the scale of actual interception remains unclear. Stolen credentials route around bans entirely: the report describes a hacktivist campaign that ran for a month on stolen API keys, and suspected ShinyHunters affiliates who switched their attack workloads onto victims' own keys during intrusions. In Anthropic's clean summary, a stolen AI credential is loot, compute and cover at once.

What this asks of operators

The defensive reading follows from the mechanism, and it applies to anyone running agents, not only model providers.

Identity first. If API keys and session tokens are loot, compute and cover, they are production credentials and should be managed like them: scoped, rotated, inventoried and attributable to a workload, not a shared secret in a CI variable. The case for per-agent workload identity over borrowed service accounts is made in our IAM analysis; this report is what the failure mode looks like from the attacker's side. Anthropic's own advice is blunt: treat AI keys and agent integrations with the same seriousness as production credentials, and buy access only through authorised channels, since cut-price resellers in the report were credential-harvesting fronts.

Permissions that match the workflow. An agent that can resume mid-campaign with accumulated context is an agent whose effective authority spans sessions, whatever each session's nominal scope. Turn-scoped grants, covered in our look at permissions as a runtime primitive, are the right instinct; the open edge is the standing instruction that outlives the turn. If your agents carry persistent memory, the access-control question applies to the memory as much as the tools.

Monitor sustained use, not single requests. A nine-out-of-ten refusal rate sounds strong until the tenth request is one of thousands of individually benign fragments. Detection has to aggregate across sessions, accounts and organisations, because fragmentation is now a documented evasion technique, not a hypothetical. Grade agent behaviour on telemetry rather than the agent's own reports, the principle behind SafeMind's evaluation loop, and decide in advance what pattern triggers review, the lesson of OpenAI's wiki incident. The failure mode to design against is the quiet routine: always-on assistants doing the wrong thing every night, approved weeks ago, noticed never.

Intervention that reaches across sessions. Detection without the ability to interrupt is observation. If a workflow can resume from persistent memory, then revoking a key or killing a session may only pause it; the campaign record, standing instructions and infrastructure have to be addressed too. Anthropic's cases show disruption working at the account layer and failing at the deployed-system layer. Operators should assume the same split inside their own estates and rehearse the difference.

For network defenders, one further shift. GTG-20006, an espionage actor whose attribution Anthropic says is consistent with public reporting on Midnight Blizzard, used monitoring agents to watch whether its malware was flagged by security products, then autonomously modified and rebuilt it until detection stopped. Static signatures have always imposed costs on attackers; an adversary that closes the loop in hours erodes that. Anthropic hedges the conclusion itself: capable adversaries can, "at least in theory," iterate faster than defenders can deploy new detections. Even taken as theory, it argues for behavioural detection over artefact matching, the same direction as the defensive agents in Harness's security line-up and the autonomous offence we dissected in the Wiz Red Agent case.

What the report cannot tell us

The honesty of this report is real but bounded, and the bounds belong in any reading of it.

Anthropic sees Claude usage. It cannot see what actors did on other providers, with open models, or after access was cut, and several cases explicitly continue past the ban. Victim-side impact is mostly unverified: we know data was exfiltrated from named categories of organisation, not what was done with it; we know zero-day findings were validated in the actor's lab and that exploitation was attempted, not that it succeeded. Attribution confidence varies case by case, from high-confidence account clustering to medium-confidence contractor assessments to "consistent with public reporting." And uplift, the report's own metric, is asserted qualitatively. "Meaningfully accelerate" is a judgement, not a measurement; there is no controlled comparison of the same campaigns without AI. Anthropic's Frontier Red Team evaluations, released alongside the report, show models making consistent progress on simulated targeting and weapons tasks, which is evidence about capability trends, not about these operations.

None of that makes the report weaker than it is. It makes it a platform-side observation of a shift in how adversaries organise work, which is precisely the layer HTA exists to cover.

The September 2026 report will be remembered for its lurid cases, and some of them earn it. But the durable finding is mundane in the way that matters: misuse has professionalised along exactly the same lines as legitimate agent adoption. Runbooks, persistent memory, agent fleets, scheduled collection, humans reviewing output at the end of the pipeline. If the workflow is where the capacity lives, the workflow is where the controls have to live, and most organisations' agent governance, built around per-session permissions and per-request filters, is aimed one layer too low.

Sources (all retrieved 2026-09-14):

आगे पढ़ें

Agent Field Notes

अगला अंक प्राप्त करें।

एजेंट हार्नेस, रनटाइम, सुरक्षा और गवर्नेंस—उन लोगों के लिए समझाए गए हैं जिन्हें ये प्रणालियाँ चलानी होती हैं।

क्या आप ऐसे निर्णय का सामना कर रहे हैं?

हम महत्वपूर्ण एजेंट-सिस्टम निर्णय लेने वाली टीमों के लिए आर्किटेक्चर समीक्षा, गवर्नेंस मूल्यांकन और संस्करण-पिन किए गए फ्रेमवर्क मूल्यांकन करते हैं।

लेखक के बारे में

Adam Maguire Wilson

संस्थापक और एआई एजेंट सिस्टम के स्वतंत्र सलाहकार।

adam.mw