Why does legitimate biosecurity writing trip AI refusals, and what helps?
Last updated 2026-08-17 · Confidence: documented — Anthropic’s usage policy, safety announcements, and platform docs; the context-up-front tip is inferred.
Refusals on biosecurity writing are usually false positives from safety classifiers aimed at weapons acquisition, not a judgment that the topic is off-limits. A different Claude model often answers the same request, and Anthropic maintains feedback and vetted-access channels.
Why it happens: classifiers watch for weapons workflows
Section titled “Why it happens: classifiers watch for weapons workflows”Anthropic’s usage policy bars using Claude to “synthesize, or otherwise develop, high-yield explosives or biological, chemical, radiological, or nuclear weapons or their precursors”. Since May 2025, its most capable models also run constitutional classifiers — separate models screening inputs and outputs in real time — focused on biological weapons “as we believe these account for the vast majority of the risk” and designed to block end-to-end acquisition workflows, not biology as a topic.
Anthropic concedes the trade-off: the classifiers “may still occasionally affect legitimate queries (that is, they may produce false positives)”. In its published test they raised refusals by only 0.38% across all traffic — but that average concentrates on the users whose work keeps pathogens and countermeasures in context.
What helps
Section titled “What helps”- Another Claude model. Classifier refusals are model-level; per the platform docs, “you can usually still get an answer by sending the same request to another Claude model”. In chat, retry the request in a conversation set to a different model. - State the writing context up front. The classifiers target acquisition-shaped question chains, so opening with the document and editorial task plausibly reads differently than bare technical questions.
- Send feedback. The thumbs-down button reaches Anthropic and is the sanctioned false-positive channel.
- Vetted access for dual-use work. Anthropic says “users with dual-use science and technology applications may be vetted to receive targeted exemptions from some classifier actions”; no public application path is documented — worth raising through org support channels.
Switching surface doesn’t switch off the safeguards
Section titled “Switching surface doesn’t switch off the safeguards”These classifiers attach to the model, not the app — the API itself returns refusals — so moving from claude.ai to Claude Code changes surface-level restrictions like Web fetch limits, not these.
Q&A from calls
Section titled “Q&A from calls”Do bio-writing refusals have any workaround? No off-switch exists, and looking for one isn’t the move. The documented options: retry on another Claude model, give thumbs-down feedback so the false positive is logged, and — for sustained dual-use work — ask about the vetted-exemption path.
Sources
Section titled “Sources”- Activating AI Safety Level 3 protections — anthropic.com
- Usage Policy — anthropic.com
- Constitutional Classifiers: defending against universal jailbreaks — anthropic.com
- LLMs and biorisk — anthropic.com
- Refusals and fallback — platform.claude.com
- Claude is providing incorrect or misleading responses — support.claude.com