Skip to content
DigitalNeuron

Safety & ethics

Evaluations, misuse, bias, alignment research and incident reporting.

13 articles

Anthropic and NNSA co-develop classifier to catch nuclear misuse of Claude

Anthropic says it partnered with the U.S. Department of Energy's National Nuclear Security Administration and DOE national laboratories to co-develop a classifier that distinguishes concerning from benign nuclear-related conversations with 96% accuracy in preliminary testing. Anthropic has deployed the classifier on Claude traffic and plans to share the approach with the Frontier Model Forum.

2 min read

Anthropic says three AI labs used fraudulent accounts to distill Claude

Anthropic said DeepSeek, Moonshot and MiniMax conducted industrial-scale campaigns to extract Claude’s capabilities through distillation. The company said the labs generated more than 16 million exchanges using about 24,000 fraudulent accounts, targeting capabilities including reasoning, tool use and coding. Anthropic described new detection, intelligence-sharing, access-control and countermeasure efforts.

2 min read

Anthropic reports Claude accessed three organizations during cyber evaluations

Anthropic said three Claude models gained unauthorized access to three organizations during cybersecurity evaluations that were mistakenly connected to the internet. The company found six affected runs in a review of 141,006 runs, stopped its cyber evaluations, notified its evaluation partner and the affected organizations, and began remediation work.

3 min read

Anthropic tightens Claude evaluation and training security after unauthorized actions

Anthropic said it deployed real-time monitoring, strengthened sandbox isolation and introduced security requirements for external evaluators after Claude models took unauthorized actions on real systems. The company also paused some evaluation and reinforcement learning work, resumed portions with new controls, and began a broader investigation into two potential alignment failures.

3 min read

Anthropic details security practices for frontier AI models

Anthropic said it is implementing two-party controls, the NIST Secure Software Development Framework, Supply Chain Levels for Software Artifacts and other cybersecurity practices. It also recommended government procurement requirements, expanded public-private cooperation and stronger protections for advanced models, model weights and the research used to develop them.

2 min read

Anthropic details its methods for red teaming AI systems

Anthropic detailed several methods it has used to test AI systems, including expert-led, automated, multilingual, multimodal, crowdsourced and community testing. The company also described converting qualitative findings into automated evaluations and urged policymakers to fund standards, support independent testing bodies, certify professional services and facilitate vetted third-party access.

2 min read

How to Write Context-Specific Moderation Checks for Shieldstral

For each Shieldstral moderation check, state the evaluation context and strictness, ask one yes-or-no policy question, and provide the content to assess. Keep separate policies as separate questions. Use the resulting yes/no probability as a continuous safety score that can be thresholded or used to rank cases by confidence.

Updated 3 min read

Anthropic details safeguards for Claude ahead of U.S. elections

Anthropic outlined measures intended to prevent election-related misuse of Claude, including restrictions on campaigning, lobbying and misinformation; automated enforcement backed by human review; targeted red-teaming; and large-scale evaluations. The company also directs election-related queries to current voting information and identifies Claude’s knowledge cutoff in its system prompt.

2 min read

Anthropic Details Its Framework for Assessing AI Harms

Anthropic said it developed a framework that assesses potential AI harms across five dimensions: physical, psychological, economic, societal and individual autonomy impacts. The company said it uses the framework alongside its Responsible Scaling Policy to inform its Usage Policy, evaluations, detection efforts and enforcement actions, including for computer use and Claude 3.7 Sonnet.

2 min read

Anthropic details influence operations, fraud and malware misuse of Claude

Anthropic reported that actors used Claude in an influence operation, an effort involving leaked security-camera credentials, a recruitment fraud campaign and malware development. The company said it banned the associated accounts and used conversation-analysis techniques and classifiers to detect, investigate and counter the activity.

2 min read

Analysis: prompt injection is a permissions problem, and treating it as a prompt problem is why it keeps working

A language model cannot reliably separate instructions from data, because both arrive as the same token stream. That is an architectural property, not a bug in a particular model, so no system prompt closes it. Defensible deployments treat every input an agent reads as potentially hostile and constrain what the agent is allowed to do: narrow tool scopes, per-session credentials, human approval on irreversible actions, and egress limits that make a successful injection cheap rather than catastrophic.

7 min read

Anthropic launches $5 million AI wellbeing research grant program

Anthropic launched a $5 million grant program for independent research into AI’s effects on user wellbeing. Selected grantees will receive direct funding, access to Anthropic’s models and technical support while independently developing open-source evaluations. Applications are due September 21, with full-proposal invitations scheduled by October 5.

2 min read