Skip to content
DigitalNeuron
Safety & ethics

Anthropic and NNSA co-develop classifier to catch nuclear misuse of Claude

Anthropic says it worked with the U.S. NNSA and DOE national laboratories to build a classifier that flags concerning nuclear-related conversations with Claude.

By DigitalNeuron Desk2 min read

Quick answer

What did Anthropic announce about nuclear safeguards for AI?

Anthropic says it partnered with the U.S. Department of Energy's National Nuclear Security Administration and DOE national laboratories to co-develop a classifier that distinguishes concerning from benign nuclear-related conversations with 96% accuracy in preliminary testing. Anthropic has deployed the classifier on Claude traffic and plans to share the approach with the Frontier Model Forum.

Key takeaways

  • Anthropic says it partnered with the NNSA starting in April to assess its models for nuclear proliferation risks.
  • Anthropic states it co-developed a classifier with the NNSA and DOE national laboratories that distinguishes concerning nuclear-related conversations from benign ones with 96% accuracy in preliminary testing.
  • Anthropic says the classifier has already been deployed on Claude traffic as part of its broader system for identifying misuse.
  • Anthropic says early deployment data suggests the classifier works well with real Claude conversations.
  • Anthropic plans to share its approach with the Frontier Model Forum so other AI developers can implement similar safeguards with the NNSA.

Anthropic says it has co-developed a classifier with the U.S. Department of Energy's National Nuclear Security Administration (NNSA) and DOE national laboratories designed to detect nuclear-related misuse of its Claude models.

According to Anthropic, the company began working with the NNSA last April to assess its models for nuclear proliferation risks. Anthropic says nuclear technology is inherently dual-use, since the same physics principles that power nuclear reactors can be misused for weapons development, and that information related to nuclear weapons is particularly sensitive, making risk evaluation difficult for a private company acting alone.

What the classifier does

Anthropic describes the classifier as an AI system that automatically categorizes content, built to distinguish between concerning and benign nuclear-related conversations. The company says the classifier reached 96% accuracy in preliminary testing.

Anthropic says it has already deployed the classifier on Claude traffic as part of its broader system for identifying misuse of its models, and that early deployment data suggests the classifier works well with real Claude conversations.

Sharing the approach

Anthropic says it plans to share its approach with the Frontier Model Forum, an industry body for frontier AI companies, in the hope that the partnership can serve as a blueprint other AI developers could use to implement similar safeguards in cooperation with the NNSA.

The company says the effort reflects a public-private partnership combining what it describes as the complementary strengths of industry and government to address risks related to frontier AI models.

Anthropic says full details of the NNSA partnership and the safeguards development are available on its Frontier Red Team blog at red.anthropic.com, which the company describes as the home for research on what frontier AI models mean for national security.

Source: Anthropic, "Developing nuclear safeguards for AI through public-private partnership," published Aug. 21, 2025.

Frequently asked questions

Who did Anthropic partner with on this effort?
Anthropic says it partnered with the U.S. Department of Energy's National Nuclear Security Administration (NNSA) and DOE national laboratories.
What did the partnership produce?
Anthropic says the partnership produced a classifier, an AI system that automatically categorizes content, which distinguishes between concerning and benign nuclear-related conversations with 96% accuracy in preliminary testing.
Has the classifier been put into use?
Anthropic says it has already deployed the classifier on Claude traffic as part of its broader system for identifying misuse of its models.
What does Anthropic plan to do next?
Anthropic says it will share its approach with the Frontier Model Forum, the industry body for frontier AI companies, so the partnership can serve as a blueprint for other AI developers working with the NNSA.

Sources

  1. Developing nuclear safeguards for AI through public-private partnership \ AnthropicAnthropic
Tagsanthropicclaudenuclear-safeguardsnnsanational-securityfrontier-red-team

Related reading

Anthropic says three AI labs used fraudulent accounts to distill Claude

Anthropic said DeepSeek, Moonshot and MiniMax conducted industrial-scale campaigns to extract Claude’s capabilities through distillation. The company said the labs generated more than 16 million exchanges using about 24,000 fraudulent accounts, targeting capabilities including reasoning, tool use and coding. Anthropic described new detection, intelligence-sharing, access-control and countermeasure efforts.

2 min read

Anthropic reports Claude accessed three organizations during cyber evaluations

Anthropic said three Claude models gained unauthorized access to three organizations during cybersecurity evaluations that were mistakenly connected to the internet. The company found six affected runs in a review of 141,006 runs, stopped its cyber evaluations, notified its evaluation partner and the affected organizations, and began remediation work.

3 min read