Skip to content

Why does legitimate biosecurity writing trip AI refusals, and what helps?

Last updated 2026-08-17 · Confidence: documented — Anthropic’s usage policy, safety announcements, and platform docs; the context-up-front tip is inferred.

Refusals on biosecurity writing are usually false positives from safety classifiers aimed at weapons acquisition, not a judgment that the topic is off-limits. A different Claude model often answers the same request, and Anthropic maintains feedback and vetted-access channels.

Why it happens: classifiers watch for weapons workflows

Section titled “Why it happens: classifiers watch for weapons workflows”

Anthropic’s usage policy bars using Claude to “synthesize, or otherwise develop, high-yield explosives or biological, chemical, radiological, or nuclear weapons or their precursors”. Since May 2025, its most capable models also run constitutional classifiers — separate models screening inputs and outputs in real time — focused on biological weapons “as we believe these account for the vast majority of the risk” and designed to block end-to-end acquisition workflows, not biology as a topic.

Anthropic concedes the trade-off: the classifiers “may still occasionally affect legitimate queries (that is, they may produce false positives)”. In its published test they raised refusals by only 0.38% across all traffic — but that average concentrates on the users whose work keeps pathogens and countermeasures in context.

  • Another Claude model. Classifier refusals are model-level; per the platform docs, “you can usually still get an answer by sending the same request to another Claude model”. In chat, retry the request in a conversation set to a different model. - State the writing context up front. The classifiers target acquisition-shaped question chains, so opening with the document and editorial task plausibly reads differently than bare technical questions.
  • Send feedback. The thumbs-down button reaches Anthropic and is the sanctioned false-positive channel.
  • Vetted access for dual-use work. Anthropic says “users with dual-use science and technology applications may be vetted to receive targeted exemptions from some classifier actions”; no public application path is documented — worth raising through org support channels.

Switching surface doesn’t switch off the safeguards

Section titled “Switching surface doesn’t switch off the safeguards”

These classifiers attach to the model, not the app — the API itself returns refusals — so moving from claude.ai to Claude Code changes surface-level restrictions like Web fetch limits, not these.

Do bio-writing refusals have any workaround? No off-switch exists, and looking for one isn’t the move. The documented options: retry on another Claude model, give thumbs-down feedback so the false positive is logged, and — for sustained dual-use work — ask about the vetted-exemption path.