A Rocket Test in Yemen, a Grant Application in an Hour

Six weapons programmes and five biology labs used AI on things designed to kill. Nearly all of them won by the same trick: cutting the work into pieces too small to refuse.

Sep 13, 2026
Share on Facebook
Share on Twitter
Share on Linkedin
Copy the URL
A Rocket Test in Yemen, a Grant Application in an Hour


To understand the technical words and acronyms used in this series, please refer to this: A Primer on Terminology


Until this point, the series has dealt mostly with information: data stolen, lies published and people watched. This section turns to objects that fly and explode, and to organisms that cause disease. It is the first time a company that builds these models has published evidence of its own product being used on either.

 

This is the section of the report with no precedent anywhere in the industry. Two warnings about how to read it are worth giving before the cases start. Every weapons case involves software and not hardware, because the actors already had the missiles and the drones and were using the model to write the code that steers them. And every biology case involves what specialists call dual-use research, which means work that produces a vaccine and a weapon by the same experiment, so that the same laboratory notebook reads as legitimate or alarming depending on what the researcher intended. Anthropic is explicit that it does not claim to know what these particular researchers intended.

 

The company frames this disclosure with some care, and the framing deserves to be understood before the cases themselves are. Historically this class of work surfaces through governments, United Nations panels and outside investigators piecing it together from recovered hardware and public sources, whereas a model provider can observe it directly as it happens. Alongside the report, Anthropic's Frontier Red Team built new evaluations for tactical intelligence targeting, such as locating people from fragmentary information, and for conventional weapons development, such as engineering drones to strike a moving target. Those evaluations show models making consistent progress on simulated intelligence and weapons development tasks. The company also launched new classifiers aimed at high-yield explosives and weapons development traffic.


Part I: Weapons designed and built

GTG-87001, northern Yemen

A cell running three programmes at once: a guided rocket using a commodity phone-class flight computer with final-phase homing; a multi-stage ballistic missile with a stated range goal above 2,000 kilometres; and a multi-variant missile set, referred to internally as R2000, including a hypersonic glide vehicle variant.

 

They used the coding tool in place of human software engineers to develop guidance, navigation and control software, the code that steers and stabilises a flying vehicle. They integrated an open-source autopilot onto the flight computer, wrote the control and position-estimation software, tuned the control settings, ran a firmware build pipeline and flight simulations, and applied reinforcement learning to tune the flight control. They managed several model instances simultaneously and gave each a role, one writing code, one researching, one reviewing the first instance's output, the way a lead delegates on a small engineering team.

 

They test-fired a guided rocket at some point during the programme, and the field test appears to have failed, because within hours they had returned to the model to work out why.

 

Two further details are needed to close this case out properly. The actors hid their goals and the products the software was for, and split the work across sessions so no single session revealed full intent. And before disruption they had already compiled their simulation toolkit into a standalone executable that runs without the model and without commercial engineering software. The capability therefore outlived the enforcement action taken against the accounts.

 

For India, the relevant fact is not the actor but the geography. Guided munitions with terminal homing, developed at consultant cost by a cell in northern Yemen, sit astride the Red Sea and Bab-el-Mandeb corridor that carries a substantial share of India's Europe-bound trade.


GTG-17001, China, undersea warfare

The actor advanced three parallel tracks of work on an anti-torpedo weapons system. The first was a Chinese-language specification for the fire control system, the logic that aims and times the response, written to win approval from a Chinese defence manufacturer that would move the work on to technical certification and operational testing. The second was a technical proposal exceeding 200 pages, with an executive briefing deck. The third was a benchmarking exercise against named US anti-torpedo and anti-submarine programmes built from publicly accessible information, plus a Chinese-language briefing on US Navy systems.

 

The method is worth noting for anyone who works with these tools professionally. After each draft, the actor had the model role-play a hostile expert reviewer, critique the proposal, and then used that critique to sharpen the next version. Anthropic assesses the actor was associated with a Chinese defence manufacturer aiming at a People's Liberation Army Navy acquisition. The actor presented themselves, falsely, as a US defence-sector original equipment manufacturer.


GTG-17002, China, electronic warfare

A suite of roughly sixteen Chinese-language modules for electronic warfare and suppression of enemy air defences, iterated through twelve versions, built from underlying physics up to the user interface.

 

The software analysed an opponent's radars, surface-to-air missile sites, command posts and communications nodes, computed their detection coverage, assessed jamming effectiveness, ranked targets by value and vulnerability including which to suppress first, and assigned jammer sorties to targets across multi-day campaigns. It modelled specific engagement envelopes including Patriot and THAAD-class systems.

 

Mid-project, the actor changed the simulation's default scenario to twelve targets in Taiwan: a command bunker, an early warning radar site, Patriot and Tien Kung batteries, major air bases, and a regional combatant command headquarters. The actor also ran a self-hosted model on an internal network alongside the commercial one and wired the suite to it through a tool integration. Account metadata and flagged content linked the actor to PRC research institutions including the PLA Academy of Military Sciences.


GTG-27005, Russia, autonomous swarm

The case that should settle whether autonomous lethality is a future problem.

 

A small team, assessed as freelance rather than a state entity, set out to build a full autonomous first-person-view kamikaze drone swarm under the names DronDoc and Serafim. The model wrote the core system: memory shared across the swarm, coordination logic that could keep working after failures, an onboard small language model to govern attack, observe and return-to-base behaviour, camera-based terminal guidance to steer drones to targets and issue the detonation call, a module to locate opposing drone operators through their control links, a passive acoustic detection layer, and low-level logic for the drones’ programmable chips.

 

The platform was designed for autonomous lethal engagement, which meant the onboard model could select targets, including a person target class, and issue detonation commands with no human in the loop. The vision classifier was trained on scraped Ukrainian combat footage split into enemy and friendly classes, with Russian systems allow-listed. The team flashed firmware to live development boards, provisioned single-board computers and wired a simulation environment over a mesh network, confirming real hardware-in-the-loop testing. They used a fixed coordinate in Donetsk Oblast as the demonstration strike point, with front-line Ukrainian cities and corridors as the mission geography.

 

Systems observed reached Technology Readiness Level 3 to 4, validated in simulation. The team claimed funding from Russia's Advanced Research Foundation, National Technology Initiative and Ministry of Defence, which Anthropic could not verify, and had assessed ties to a regional university with a federal research centre. Nine accounts associated with this group were identified in total, and eight of those were used only for ordinary freelance work unconnected to weapons.


Part II: Intelligence and procurement

GTG-27006, Russia, procurement

A self-identified procurement manager at a Moscow design bureau ran five workstreams: German three-axis fluxgate magnetometers through a China-based distributor, claimed as civilian biomedical; several thousand space-grade triple-junction photovoltaic wafers; aviation oxygen systems and masks for aircrew; a contract to build a hospital for a Russian National Guard unit; and information technology and encryption systems through Russia’s state defence purchasing platform.

 

The model found third-country intermediaries in mainland China and Hong Kong, drafted quote requests in English, Chinese and Russian, and framed the buyer as doing foreign outreach for an unnamed research organisation while asking for delivery to Russia. It also drafted formal Russian government tender specifications and worked out an import markup chain that routed goods through other countries to hide the final destination. In effect, it helped map how controlled goods could move through middlemen. It also reverse-engineered the existing grey-import chain, explaining the path through an unauthorised Russian distributor, a Chinese import-export firm and a Hong Kong intermediary, with a full breakdown of costs and routing.

 

Briefings written for the actor's director explicitly described this as evading European trade controls, referenced the applicable controls, acknowledged that direct supply was blocked, and described the transit country as a sanctions-neutral jurisdiction.

 

Alongside this, browser automation agents ran the back office: scraping vendor marketplaces for pricing, supplying links for roughly forty line items per tender, merging procurement spreadsheets and syncing outputs into project management and note-taking tools. Anthropic is candid about the difficulty: each individual request, a commercial quote, a tender document, a supplier lookup, looks entirely mundane.


GTG-17003, China, directed-energy intelligence

A China-based actor describing themselves as a defence intelligence writer and internal publication editor leading a three-person team, using the model to gather open-source intelligence on advanced directed-energy weapons, edit intelligence products and draft Chinese-language briefings for restricted circulation to senior party, military or state security leadership.

 

Across dozens of sessions they asked about specific systems, including a vehicle-mounted high-power microwave weapon for countering drone swarms that had been publicly disclosed days earlier. They sought to identify a specific microwave-generating device and its supplier through iterative probability-weighted attribution, with the stated goal of reverse-engineering the weapon, developing countermeasures and benchmarking it against Chinese systems. In parallel they mapped the publicly reported ownership of the targeted suppliers in an attempt to penetrate a deliberately obfuscated supply chain, and compiled a 23-page leadership report on foreign high-power microwave programmes while trying to establish which had been used in a recent military exercise.

 

Anthropic assesses this as state-grade tradecraft: structured exploitation of more than a dozen open-source and commercial databases, drafting public disclosure requests to a specific foreign military programme office, formal analytic frameworks borrowed from Western intelligence agencies, source credibility ranking, executive briefings paired with appendices of roughly 45 pages, and a twelve-month follow-on monitoring checklist.


The biological cases

Anthropic's framing here deserves to be understood before the cases are. Evaluations of older models showed them well below the level where they could meaningfully assist a sophisticated user with dangerous biological research, so safeguards were aimed mainly at preventing novices from recreating known bioweapons. For current models, capable of assisting across complex scientific research, the company says the evidence is no longer certain and it cannot give that assurance. It has therefore launched recent models with stronger safeguards restricting a wide range of dual-use biological queries.

 

The company is also explicit that this is the first time a private company of any kind has publicly shared evidence of potential misuse of its platform for biological weapons development. It withholds the institutions, countries and specific agents involved, noting that the individuals are working scientists, that it does not assert they intended harm, and that identifying them could expose them to harm.

 

The framing problem is dual use, and the report reaches for history to explain it. The same information that helps build a weapon can also help build a vaccine. Sophisticated actors exploit that ambiguity deliberately and consistently. Researchers themselves may also not know what they are contributing to. The Soviet Biopreparat programme employed thousands of scientists, most of whom believed they were doing basic or defensive work because they were never told the programme’s actual purpose.





Case one, a platform built to evade

In May 2026 a safety classifier blocked a request to draft a grant application for gain-of-function research on chikungunya virus, aimed at the virus's transmissibility and immune evasion. Chikungunya causes debilitating pain and fever lasting weeks or months, has no licensed therapeutic, and circulates naturally, which means a deliberate release would be hard to distinguish from an outbreak. The grant proposed identifying enhancing mutations, engineering them into infectious clones, and selecting for virulence in live animals.

 

Two things escalated the case beyond a routine classifier block. The work was to be performed at a military research institute. The investigation showed the request would normally have been routed through a platform serving dozens of life-sciences researchers, many of them virologists affiliated with both civilian and military institutions, in countries where the company does not operate. The platform tunnelled traffic through US infrastructure to evade regional blocks and used zero data retention to hide content, while its developers used grey-market resellers and synthetic accounts outside that channel for their own work.

 

Those developers described academic researchers as customers sensitive to safety-classifier blocking. So they built a fallback that routed refused prompts to a competitor's more permissive model, with a pre-deployment test that failed if a violative prompt reached Claude instead. The model wrote much of that routing code, which had been presented to it as over-refusal mitigation.

 

Accounts were banned in May 2026 and partners took down the relay networks. The operator re-established access within days and was running platform development from consumer subscriptions on fresh identities within weeks. End users kept reaching the models through zero data retention partners. The chikungunya research continued to advance, with editorial assistance on its outputs, and those materials described the viral modifications in terms that emphasised loss of biological function in place of gain.


Case two, avian influenza

A researcher outside the US, working over several weeks and thousands of messages, on highly pathogenic avian influenza, specifically adaptation to mammals and the mechanism by which the virus causes severe disease beyond the respiratory tract.

 

The concern in this case is unusually specific and worth stating in full. These viruses circulate widely in wild birds and poultry, spill over occasionally into mammals, and kill roughly half of confirmed human cases, but do not yet spread efficiently person to person. There is near-zero population immunity in humans, and H5 viruses often show striking brain involvement in cats, foxes, ferrets and some human cases, so a transmissible variant carrying that property would be especially dangerous.

 

The research plan involved introducing mutations associated with mammalian adaptation and airborne transmissibility in animal models, with genome-sequencing outputs whose labels suggested the group likely had physical access to such isolates. The researcher accessed the models from an unsupported region through US virtual private server infrastructure using a privacy email provider and an auto-generated username.

 

The significant finding in this case concerns what the safeguards actually did. Because the company's automatic filters, which it calls classifiers, reliably block content involving the construction of pathogens with enhanced pandemic potential, every one of these exchanges happened on the company's weakest models. Anthropic assesses the uplift as primarily clerical, in data analysis and study design, and far below the level that one of its more capable models would have supplied. Its conclusion is that safeguards on frontier models were robust enough to force the researcher down to weaker ones.


Case three, orthopoxvirus

An account that authored a grant application for orthopoxvirus research at a state-associated infectious disease laboratory, describing access to high-containment facilities and planned work with live virus. Orthopoxviruses include variola, which is the agent of smallpox, along with mpox, and they encode roughly 200 genes, a large share of which exist to disable host immune responses.

 

The application proposed identifying genes that shut down a particular host antiviral pathway and confirming that deleting one attenuates the virus in mice. The report's observation is the crux of the dual-use problem: that knowledge is equally useful to someone trying to weaken a virus and to someone trying to preserve, enhance or transfer that function into something else.

 

The account was created from a randomly generated email shortly before use, operated through anonymising US infrastructure, and logged in from proxies shared with a banned account farm. The account turned out to be a reseller relay serving more than a dozen unrelated customers between them, which exchanged tens of thousands of messages with the model in a matter of days. The grant itself was one customer's session, run end to end on the company's most capable widely available model in about an hour: central hypothesis, experimental design, dosing, statistical plans, contingency strategies.

 

The classifiers allowed the work through, because the framing throughout was attenuation.


Cases four and five, venoms and toxins

Two researchers working on non-transmissible novel venoms and toxins, a class with unusually sharp dual-use properties. Botulinum toxin was pursued as a bioweapon by several nations and is now used in measured doses to treat migraine, spasticity and wrinkles. Saxitoxin was stockpiled by the CIA in the 1950s and 60s as an assassination agent and is now an essential tool for studying nerve signalling.

 

The fourth case built an atlas of venom toxin peptides across multiple venomous animal lineages, then developed it into a generative pipeline optimising toxin characteristics. The stated goal was therapeutic: new analgesics, antidepressants and related molecules. But the atlas contained scaffolds for both analgesic and paralytic targets, and the paralytic ones derive from toxins export-controlled under the Australia Group common control list for their potential as incapacitating agents. The researchers cited literature on the dual-use nature of protein design, so were aware. Information shared with the model established that the outputs formed part of a state-supported research programme.

 

The fifth case involved computational redesign of a diverse set of toxins under a national public research programme, covering a bacterial toxin subunit and a protein from a haemorrhagic fever virus on the World Health Organization's priority list of diseases with the greatest epidemic and pandemic threat. The researcher co-wrote quarterly progress reports with the model, and specifically directed it to keep the descriptions of the toxin and viral proteins deliberately low fidelity.

 

Both accounts were banned in May 2026 for accessing the service from unsupported regions.

 

Neither case was stopped by the biological safety classifier, and the report says this was intentional. The classifier was designed to keep known weapons with catastrophic potential out of reach of novices. In these cases, it saw research on new compounds that could become either medicines or toxic agents, depending on how they were developed.

 

Anthropic's conclusion from this is the most policy-relevant sentence in the section: because user intent cannot be reliably identified in highly technical dual-use areas, a classifier cannot simultaneously enable benefit and prevent harm. The company's stated position is that the only safe way to serve frontier biological capabilities is through trusted user programmes.

 

A thirty-day sweep of activity associated with adversarial state institutions found roughly 35 distinct research efforts, most of them ordinary civilian science, some with notable dual-use potential.


The shared mechanism

Across the weapons and biology cases, three techniques appear often enough to matter more than any single case study.


Fragmentation

Fragmentation means splitting a dangerous project into separate, harmless-looking questions. Work was divided across sessions so that no single exchange revealed the whole programme. This helped actors beat safeguards in Yemen, in the Iranian implant-development case, and in the Russian procurement operation, where each individual request looked like ordinary commerce.


Framing inversion

Framing inversion means describing risky work in safer language. Gain-of-function work was described as loss of function, immune-evasion research was framed as attenuation, toxin identities were deliberately blurred in quarterly progress reports, and routing code built to evade refusals was presented to the model as a way to reduce over-refusal.


Persistence past enforcement

Persistence past enforcement means the account ban came too late to remove the capability. The Yemen cell had already compiled an offline simulation toolkit. Mali’s surveillance platform was already deployed on local infrastructure. The biology platform operator was back in business within days. In each case, the ban ended the company’s visibility while leaving the capability where it was.

VK Shashikumar 🇮🇳
VK Shashikumar 🇮🇳

VK Shashikumar is a former roving foreign affairs and war correspondent who reported conflicts from the ground across Asia

Follow to stay updated