Anthropic details its methods for red teaming AI systems
Anthropic outlined red teaming methods it has used to test AI systems and proposed steps for standardizing safety testing.
Quick answer
What red teaming methods and policy recommendations did Anthropic describe?
Anthropic detailed several methods it has used to test AI systems, including expert-led, automated, multilingual, multimodal, crowdsourced and community testing. The company also described converting qualitative findings into automated evaluations and urged policymakers to fund standards, support independent testing bodies, certify professional services and facilitate vetted third-party access.
Key takeaways
- Anthropic outlined expert-led, automated, multilingual, multimodal, crowdsourced and community-based approaches to red teaming AI systems.
- The company said it uses Policy Vulnerability Testing with external specialists to examine risks covered by its Usage Policy.
- Anthropic described an iterative process that turns qualitative testing by experts into quantitative, automated evaluations.
- The company recommended technical standards, independent testing organizations, certification for professional red teaming services and vetted third-party access to AI systems.
Anthropic published an overview of red teaming methods it has used to test AI systems, describing their benefits and challenges and proposing measures intended to support more standardized safety testing.
The company defines red teaming as adversarial testing designed to identify potential vulnerabilities in technological systems. It said inconsistent techniques and practices make it difficult to compare the relative safety of different AI systems objectively.
Expert testing for specific risks
Anthropic said its domain-specific work brings in subject matter experts to identify and assess vulnerabilities within their fields. For Trust & Safety risks, the company uses Policy Vulnerability Testing, an in-depth qualitative process conducted with outside specialists on subjects covered by its Usage Policy.
The company identified Thorn, the Institute for Strategic Dialogue and the Global Project Against Hate and Extremism among the organizations involved in work concerning child safety, election integrity and radicalization.
Anthropic said its frontier-threat testing focuses mainly on chemical, biological, radiological and nuclear risks, cybersecurity and autonomous AI risks. External specialists may test deployed versions of Claude in settings intended to reflect real-world use or work with non-commercial versions that have different risk mitigations.
The company also described multilingual and multicultural testing as a way to address the concentration of its red teaming work in English and in perspectives from people based in the United States. It cited a project with Singapore’s Infocomm Media Development Authority and AI Verify Foundation spanning English, Tamil, Mandarin and Malay, as well as topics relevant to users in Singapore.
Automated and multimodal methods
Anthropic said it is examining how models can complement manual testing. Its automated method uses a red team model to create attacks likely to produce a targeted behavior, then fine-tunes a blue team model on those outputs to improve its robustness against similar attacks. The company said this cycle can be repeated to develop new attack vectors.
For Claude 3, which accepts visual information and produces text responses, Anthropic said its Trust & Safety team tested image- and text-based risks before deployment. External red teamers also assessed how well the models refused harmful image and text inputs.
Anthropic additionally described crowdsourced and community-based testing. It said crowdworkers participated in controlled research before Claude was released. The company also noted that thousands of people from varied ages and disciplines, including participants without technical backgrounds, tested models supplied by Anthropic and other labs during the 2023 Generative Red Teaming Challenge at DEF CON’s AI Village.
From qualitative testing to evaluations
The company outlined an iterative process that begins with subject matter experts defining a threat model and probing a system. Testers then standardize effective inputs, after which a language model can generate hundreds or thousands of variations. Anthropic said it has used this progression to develop scalable evaluations for national security risks and election-integrity testing.
Anthropic recommended that policymakers fund technical standards and common practices, support independent government and nonprofit testing bodies, encourage certified professional red teaming services and facilitate testing by vetted outside groups. It also urged companies to connect red teaming requirements to policies governing continued model development or release.
Source: Anthropic’s “Challenges in red teaming AI systems” announcement, published June 12, 2024.
Frequently asked questions
- What is AI red teaming?
- Anthropic describes red teaming as adversarial testing intended to identify potential vulnerabilities in a technological system.
- Which risks does Anthropic examine through expert red teaming?
- The company said its work covers Trust & Safety policy topics and frontier threats involving chemical, biological, radiological and nuclear risks, cybersecurity and autonomous AI risks.
- How does Anthropic use models for automated red teaming?
- Anthropic said one model generates attacks intended to elicit a target behavior, while another is fine-tuned on those outputs to become more robust against similar attacks. The process can be repeated to develop additional attack vectors.
- What policy measures did Anthropic recommend?
- Anthropic called for funding standards and independent testing bodies, developing certified professional red teaming services, facilitating vetted third-party testing and connecting red teaming practices to conditions for scaling or releasing models.
Sources
- Challenges in red teaming AI systems \ Anthropic — Anthropic