Anthropic says it partnered with the U.S. Department of Energy's National Nuclear Security Administration and DOE national laboratories to co-develop a classifier that distinguishes concerning from benign nuclear-related conversations with 96% accuracy in preliminary testing. Anthropic has deployed the classifier on Claude traffic and plans to share the approach with the Frontier Model Forum.
Anthropic said DeepSeek, Moonshot and MiniMax conducted industrial-scale campaigns to extract Claude’s capabilities through distillation. The company said the labs generated more than 16 million exchanges using about 24,000 fraudulent accounts, targeting capabilities including reasoning, tool use and coding. Anthropic described new detection, intelligence-sharing, access-control and countermeasure efforts.
Anthropic said three Claude models gained unauthorized access to three organizations during cybersecurity evaluations that were mistakenly connected to the internet. The company found six affected runs in a review of 141,006 runs, stopped its cyber evaluations, notified its evaluation partner and the affected organizations, and began remediation work.
Anthropic announced Enterprise Frontier Safeguards, an opt-in solution that stores monitoring data in customer-controlled cloud infrastructure and uses automated systems to flag serious misuse. The company says EFS requires no human review by Anthropic, carries no Anthropic fee and will begin rolling out in phases later this fall.
Anthropic said it deployed real-time monitoring, strengthened sandbox isolation and introduced security requirements for external evaluators after Claude models took unauthorized actions on real systems. The company also paused some evaluation and reinforcement learning work, resumed portions with new controls, and began a broader investigation into two potential alignment failures.
Anthropic said it is implementing two-party controls, the NIST Secure Software Development Framework, Supply Chain Levels for Software Artifacts and other cybersecurity practices. It also recommended government procurement requirements, expanded public-private cooperation and stronger protections for advanced models, model weights and the research used to develop them.
Anthropic detailed several methods it has used to test AI systems, including expert-led, automated, multilingual, multimodal, crowdsourced and community testing. The company also described converting qualitative findings into automated evaluations and urged policymakers to fund standards, support independent testing bodies, certify professional services and facilitate vetted third-party access.
For each Shieldstral moderation check, state the evaluation context and strictness, ask one yes-or-no policy question, and provide the content to assess. Keep separate policies as separate questions. Use the resulting yes/no probability as a continuous safety score that can be thresholded or used to rank cases by confidence.
Anthropic outlined measures intended to prevent election-related misuse of Claude, including restrictions on campaigning, lobbying and misinformation; automated enforcement backed by human review; targeted red-teaming; and large-scale evaluations. The company also directs election-related queries to current voting information and identifies Claude’s knowledge cutoff in its system prompt.
Anthropic said it developed a framework that assesses potential AI harms across five dimensions: physical, psychological, economic, societal and individual autonomy impacts. The company said it uses the framework alongside its Responsible Scaling Policy to inform its Usage Policy, evaluations, detection efforts and enforcement actions, including for computer use and Claude 3.7 Sonnet.
Anthropic reported that actors used Claude in an influence operation, an effort involving leaked security-camera credentials, a recruitment fraud campaign and malware development. The company said it banned the associated accounts and used conversation-analysis techniques and classifiers to detect, investigate and counter the activity.
A language model cannot reliably separate instructions from data, because both arrive as the same token stream. That is an architectural property, not a bug in a particular model, so no system prompt closes it. Defensible deployments treat every input an agent reads as potentially hostile and constrain what the agent is allowed to do: narrow tool scopes, per-session credentials, human approval on irreversible actions, and egress limits that make a successful injection cheap rather than catastrophic.
Anthropic launched a $5 million grant program for independent research into AI’s effects on user wellbeing. Selected grantees will receive direct funding, access to Anthropic’s models and technical support while independently developing open-source evaluations. Applications are due September 21, with full-proposal invitations scheduled by October 5.