Skip to content
DigitalNeuron
Safety & ethics

Anthropic details its methods for red teaming AI systems

Anthropic outlined red teaming methods it has used to test AI systems and proposed steps for standardizing safety testing.

By DigitalNeuron Desk2 min read

Quick answer

What red teaming methods and policy recommendations did Anthropic describe?

Anthropic detailed several methods it has used to test AI systems, including expert-led, automated, multilingual, multimodal, crowdsourced and community testing. The company also described converting qualitative findings into automated evaluations and urged policymakers to fund standards, support independent testing bodies, certify professional services and facilitate vetted third-party access.

Key takeaways

  • Anthropic outlined expert-led, automated, multilingual, multimodal, crowdsourced and community-based approaches to red teaming AI systems.
  • The company said it uses Policy Vulnerability Testing with external specialists to examine risks covered by its Usage Policy.
  • Anthropic described an iterative process that turns qualitative testing by experts into quantitative, automated evaluations.
  • The company recommended technical standards, independent testing organizations, certification for professional red teaming services and vetted third-party access to AI systems.

Anthropic published an overview of red teaming methods it has used to test AI systems, describing their benefits and challenges and proposing measures intended to support more standardized safety testing.

The company defines red teaming as adversarial testing designed to identify potential vulnerabilities in technological systems. It said inconsistent techniques and practices make it difficult to compare the relative safety of different AI systems objectively.

Expert testing for specific risks

Anthropic said its domain-specific work brings in subject matter experts to identify and assess vulnerabilities within their fields. For Trust & Safety risks, the company uses Policy Vulnerability Testing, an in-depth qualitative process conducted with outside specialists on subjects covered by its Usage Policy.

The company identified Thorn, the Institute for Strategic Dialogue and the Global Project Against Hate and Extremism among the organizations involved in work concerning child safety, election integrity and radicalization.

Anthropic said its frontier-threat testing focuses mainly on chemical, biological, radiological and nuclear risks, cybersecurity and autonomous AI risks. External specialists may test deployed versions of Claude in settings intended to reflect real-world use or work with non-commercial versions that have different risk mitigations.

The company also described multilingual and multicultural testing as a way to address the concentration of its red teaming work in English and in perspectives from people based in the United States. It cited a project with Singapore’s Infocomm Media Development Authority and AI Verify Foundation spanning English, Tamil, Mandarin and Malay, as well as topics relevant to users in Singapore.

Automated and multimodal methods

Anthropic said it is examining how models can complement manual testing. Its automated method uses a red team model to create attacks likely to produce a targeted behavior, then fine-tunes a blue team model on those outputs to improve its robustness against similar attacks. The company said this cycle can be repeated to develop new attack vectors.

For Claude 3, which accepts visual information and produces text responses, Anthropic said its Trust & Safety team tested image- and text-based risks before deployment. External red teamers also assessed how well the models refused harmful image and text inputs.

Anthropic additionally described crowdsourced and community-based testing. It said crowdworkers participated in controlled research before Claude was released. The company also noted that thousands of people from varied ages and disciplines, including participants without technical backgrounds, tested models supplied by Anthropic and other labs during the 2023 Generative Red Teaming Challenge at DEF CON’s AI Village.

From qualitative testing to evaluations

The company outlined an iterative process that begins with subject matter experts defining a threat model and probing a system. Testers then standardize effective inputs, after which a language model can generate hundreds or thousands of variations. Anthropic said it has used this progression to develop scalable evaluations for national security risks and election-integrity testing.

Anthropic recommended that policymakers fund technical standards and common practices, support independent government and nonprofit testing bodies, encourage certified professional red teaming services and facilitate testing by vetted outside groups. It also urged companies to connect red teaming requirements to policies governing continued model development or release.

Source: Anthropic’s “Challenges in red teaming AI systems” announcement, published June 12, 2024.

Frequently asked questions

What is AI red teaming?
Anthropic describes red teaming as adversarial testing intended to identify potential vulnerabilities in a technological system.
Which risks does Anthropic examine through expert red teaming?
The company said its work covers Trust & Safety policy topics and frontier threats involving chemical, biological, radiological and nuclear risks, cybersecurity and autonomous AI risks.
How does Anthropic use models for automated red teaming?
Anthropic said one model generates attacks intended to elicit a target behavior, while another is fine-tuned on those outputs to become more robust against similar attacks. The process can be repeated to develop additional attack vectors.
What policy measures did Anthropic recommend?
Anthropic called for funding standards and independent testing bodies, developing certified professional red teaming services, facilitating vetted third-party testing and connecting red teaming practices to conditions for scaling or releasing models.

Sources

  1. Challenges in red teaming AI systems \ AnthropicAnthropic
Tagsanthropicclaudered-teamingai-safetymodel-evaluationspolicy

Related reading

Anthropic details safeguards for Claude ahead of U.S. elections

Anthropic outlined measures intended to prevent election-related misuse of Claude, including restrictions on campaigning, lobbying and misinformation; automated enforcement backed by human review; targeted red-teaming; and large-scale evaluations. The company also directs election-related queries to current voting information and identifies Claude’s knowledge cutoff in its system prompt.

2 min read

Anthropic Details Its Framework for Assessing AI Harms

Anthropic said it developed a framework that assesses potential AI harms across five dimensions: physical, psychological, economic, societal and individual autonomy impacts. The company said it uses the framework alongside its Responsible Scaling Policy to inform its Usage Policy, evaluations, detection efforts and enforcement actions, including for computer use and Claude 3.7 Sonnet.

2 min read

Anthropic details influence operations, fraud and malware misuse of Claude

Anthropic reported that actors used Claude in an influence operation, an effort involving leaked security-camera credentials, a recruitment fraud campaign and malware development. The company said it banned the associated accounts and used conversation-analysis techniques and classifiers to detect, investigate and counter the activity.

2 min read