In a report published on September 10, Anthropic — the U.S.-based company that builds and provides the Claude AI model — said it blocked users from using the model for a range of potentially harmful activities across cybersecurity, surveillance, and biotechnology between December 2025 and August 2026.
While Anthropic has in general been more open about how users around the world are using its models, the report leaves undiscussed an older, more general problem that also raises questions about what Claude could really have done, and where Anthropic might itself be caught in a bind.
Sensing danger
Today, AI companies like Google, Anthropic, and OpenAI are mostly policing themselves while also building out an unequivocally powerful technology. In this context, self-policing should stoke some concerns: for example, that employees who leave these companies proclaiming insufficient attention to safety, like Jacob Coxon did, are probably being more theatrical than constructive; that AI companies could be talking up how beneficial their products can be while continuing to under-protect society against the damage it could wreak from the wrong hands; that public discourse around viruses-that-can-be-weaponised deserves to be had with care, since the field only recently became politically radioactive; and that developers of AI models should not release a new capability before they have also engineered adequate safeguards.
(On the last point, in fact, Anthropic CEO Dario Amodei wrote on his blog on September 12, “As models build future models, the rate of improvement may become staggeringly fast. Slowing the rate from ‘extremely fast’ to ‘only somewhat fast’ gives up relatively little strategic advantage, while potentially greatly improving safety.”)
Anthropic’s report is also candid: that its earlier Opus 4 and Sonnet 4.5 models had “less stringent” safeguards against use in biology because internal evaluations had concluded they were not capable enough, and that Anthropic decided to control access to queries about dual-use biology research only with its newer models, especially Fable 5. The company had also said in August that its filters to block conversations about biological weapons had been inactive on roughly 133 million exchanges for nearly a year.
The report has case studies of users who used the older models for biology research for many weeks, including drafting a grant proposal for gain-of-function research on the chikungunya virus. (This is where, for example, researchers augment a virus with a function it did not naturally evolve and study how that affects the virus’s evolution). Another researcher planned experiments involving avian influenza and mammals. And so on, in the life sciences domain. Anthropic also said Claude was involved in a guided rocket programme where developers conducted a field test with Claude-made code. A different developer had Claude write code for a drone swarm and tested it in a simulation. And so on. In all cases, Anthropic said it terminated the users’ accounts as soon as the potential for illicit use became clear.
Sensing danger is not always straightforward. When Enrico Fermi’s team tested the world’s first nuclear reactor, a team member named Leona Woods asked him, “When do we become scared?” It turns out machines also have a tough time answering this question.
AI models require sophisticated hardware and software setups, including software components called input and output classifiers, which respectively check the user’s inputs and the model’s outputs. As the report put it, a classifier “cannot simultaneously enable benefit and prevent harm” in contexts where some knowledge can be used for good or for bad. This means a user using Claude to design something dangerous, like guidance code for a missile, could have already saved that output offline before Anthropic spotted the misuse and closed the user’s account.
More broadly, a model can generate and show a part of its response to the user before it checks for dangerous content or triggers an enforcement action, and Anthropic can ban a user only after the relevant activity has occurred.
The Meta Ray-Ban smart glasses give Meta an opportunity — speaking purely technologically — to scan a photo captured by the glasses, check for illicit information, and decide whether the photo can be uploaded to its servers. As it happens, AI providers are actually better placed to achieve this than smart wearables or even social media platforms. Input classifiers can scan a prompt before generation begins and output classifiers inspect the response. When a model’s output exists on the provider’s servers as it is being generated and before it is streamed to the user, the provider will have an opportunity to check it before or even during delivery.
The thorn here is the potential for surveillance. In the U.S., where most of the world’s most popular AI models are being conceived, the Fourth Amendment reins in the state, but not private companies, so a company can inspect content on its own servers to enforce its terms of service.
Rifle or wheelchair?
If the means and the opportunity both exist, wouldn’t an AI company’s claim to “have safeguards” become misleading when data can still be exfiltrated before the company closes an account?
This is a useful question because it reveals two details that are often conflated: the company needs to have a checkpoint and the checkpoint needs to be able to recognise that some information presented to it is harmful. So while there is a built-in gap between output generation and output delivery, the checkpoint also needs to be able to tell whether the output is harmful — and it would be reckless to attempt to do so based on single messages.
Say you are watching a worker hand small pieces of metal out of a window to someone on the other side, and your job is to say if the someone is building a rifle or a wheelchair. At first you will not be able to say much because all you see are pins, springs, plates, and screws going through. But over time, the sequence of parts and the larger parts will reveal if the intended object is a rifle or a wheelchair. Similarly, a user may prompt Claude to write code for a control loop, which appears in guided missiles as well as in air-conditioners, refrigerators, and toilet tanks. The input and output classifiers are only on alert when a sufficient pattern has accumulated.
However, the limits of what can be inferred from a single message do not excuse AI companies as the limits also reveal what the companies’ “safeguards” actually entail. A statement that “there are safeguards” can be true to the extent that the input/output classifiers act on potentially harmful information when they detect it but false to the extent that the “safeguards” have prevented harm altogether, thanks to data exfiltration before termination.
Export controls
This is not a new problem. Older systems have dealt with similar issues for decades. But the curious fact is they did not solve the problem. Instead, they roughly converged on two alternatives, by ensuring which they were able to minimise harm of one kind: restricting the people who could access a technology, then tracking patterns of use over time. (This is also the “trusted access programmes” paradigm the Anthropic report mentions later.)
A well-known system with these features is export controls. Neither the Cold-War-era Coordinating Committee for Multilateral Export Controls nor the Wassenaar Arrangement and the Nuclear Suppliers Group of today have inspected individual shipments. Instead, to meet their conditions, participating actors have generally required an export license linked to a specific end-user and an end-use certificate, and have had to steer clear of a list of barred organisations and countries. Without having to track individual shipments of some parts or expecting them to reveal if they’re headed for a power plant or a weapons facility, thus, the regime can simply control who can access those parts at all.
Similarly, in 2011 and 2012, researchers from Erasmus Medical Centre, the University of Wisconsin-Madison, and the University of Tokyo submitted papers about engineering the H5N1 avian influenza virus to be transmissible between mammals through the air. The U.S. National Science Advisory Board for Biosecurity recommended the authors redact their methods before publishing, setting off a months-long spat between governments and journals. The journals eventually published the papers almost in full because the scientific community argued that suppressing the methods wouldn’t stop a determined state actor from recreating them but would likely slow down defensive research in other quarters.
Similar events have played out in anti-money-laundering efforts (where banks can monitor single transactions), large ammonium nitrate purchases following the 1995 Oklahoma City bombing, and sales of pseudoephedrine during the American methamphetamine crisis. For different reasons, empowering the classifier has not been the answer.
However, that still leaves the question of completely restricting access to dangerous AI models — operating with the prospect of ‘data exfiltration before termination’ — open. Anthropic has also not publicly examined how much of the extent to which it gatekeeps users today is about safety per se versus about preserving a market.
mukunth.v@thehindu.co.in


