Skip to content
DigitalNeuron
Safety & ethics

Anthropic details security practices for frontier AI models

Anthropic outlined cybersecurity controls it is implementing and recommended procurement rules and public-private cooperation to protect frontier AI models.

By DigitalNeuron Desk2 min read

Quick answer

What security measures did Anthropic announce for developing and deploying frontier AI models?

Anthropic said it is implementing two-party controls, the NIST Secure Software Development Framework, Supply Chain Levels for Software Artifacts and other cybersecurity practices. It also recommended government procurement requirements, expanded public-private cooperation and stronger protections for advanced models, model weights and the research used to develop them.

Key takeaways

  • Anthropic said it is implementing two-party controls, SSDF, SLSA and other cybersecurity practices.
  • The company recommended multi-party authorization across systems used to develop, train, host and deploy frontier AI models.
  • Anthropic encouraged extending NIST’s Secure Software Development Framework to cover model development.
  • The company said governments could initially establish the proposed controls through procurement requirements for AI companies and cloud providers.
  • Anthropic recommended treating frontier AI laboratories similarly to critical-infrastructure companies for public-private cooperation and information sharing.

Anthropic said it is implementing a set of cybersecurity practices intended to protect frontier artificial intelligence models, their weights and the research used to develop them. The company also proposed government procurement requirements and closer cooperation between AI laboratories and public agencies.

The measures described by Anthropic include two-party controls, the NIST Secure Software Development Framework, known as SSDF, and Supply Chain Levels for Software Artifacts, known as SLSA. Anthropic said stronger protections will be needed as model capabilities increase and described the work as an iterative process involving government and industry.

Multi-party authorization

Anthropic recommended applying two-party control to systems involved in developing, training, hosting and deploying frontier AI models. It calls this approach “multi-party authorization to AI-critical infrastructure design.”

Under the design Anthropic described, no individual would have persistent access to production-critical environments. A person seeking access would need a coworker’s authorization, a business justification and a time-limited permission. The company said emerging AI laboratories can implement such controls even without the resources of large enterprises.

Anthropic presented the measure as protection against advanced threat actors and insider risk. It said two-party control depends on a broader range of cybersecurity practices for correct implementation.

Secure model development

Anthropic also recommended applying secure software development practices throughout frontier model environments. It identified SSDF and SLSA as standards that can be translated to model development and the software connected to it.

The company said using the two frameworks together can create a chain of custody for a deployed AI system. Anthropic described this as a way to connect a deployed model to the company that developed it and establish provenance.

Anthropic calls this proposed combination a “secure model development framework.” It encouraged NIST to extend SSDF through its standard-setting process so that the framework explicitly encompasses model development.

Procurement and regulation

Anthropic said governments could establish multi-party authorization and secure model development as procurement requirements for AI companies and cloud providers seeking public contracts. These measures would sit alongside other cybersecurity requirements applying to the companies.

The company said procurement rules affecting U.S. cloud providers could have an effect similar to broad market regulation because many frontier model companies use their infrastructure. Anthropic proposed procurement as a step that could operate before regulatory requirements.

Anthropic said many security measures could begin as voluntary arrangements. It added that governments may eventually determine that procurement or regulatory powers should mandate compliance.

Public-private cooperation

Anthropic recommended that frontier AI research laboratories participate in public-private cooperation in the same manner as companies in critical-infrastructure sectors such as financial services. It suggested that frontier AI could be designated as a special subsector of the existing information technology sector.

According to Anthropic, such a designation could support greater information sharing and cooperation among AI laboratories and government agencies. The company said this arrangement could help laboratories guard against well-resourced malicious cyber actors.

Source: Anthropic’s “Frontier model security” announcement, published July 25, 2023.

Frequently asked questions

What is Anthropic implementing?
Anthropic said it is implementing two-party controls, the NIST Secure Software Development Framework, Supply Chain Levels for Software Artifacts and other cybersecurity practices.
What does Anthropic mean by two-party control?
Anthropic described a design in which no individual has persistent access to production-critical environments. A person must request time-limited access from a coworker and provide a business justification.
What government measures did Anthropic recommend?
Anthropic proposed procurement requirements covering AI companies and cloud providers that contract with governments. It said voluntary arrangements could come first, with procurement or regulatory mandates potentially following over time.
How does Anthropic propose organizing public-private cooperation?
Anthropic said frontier AI laboratories should cooperate with governments in a manner similar to critical-infrastructure sectors. It suggested that frontier AI could become a special subsector of the existing information technology sector.

Sources

  1. Frontier model security \ AnthropicAnthropic
Tagsanthropicclaudeai-securitycybersecurityfrontier-modelsai-policy

Related reading

Anthropic details influence operations, fraud and malware misuse of Claude

Anthropic reported that actors used Claude in an influence operation, an effort involving leaked security-camera credentials, a recruitment fraud campaign and malware development. The company said it banned the associated accounts and used conversation-analysis techniques and classifiers to detect, investigate and counter the activity.

2 min read

Anthropic details its methods for red teaming AI systems

Anthropic detailed several methods it has used to test AI systems, including expert-led, automated, multilingual, multimodal, crowdsourced and community testing. The company also described converting qualitative findings into automated evaluations and urged policymakers to fund standards, support independent testing bodies, certify professional services and facilitate vetted third-party access.

2 min read

Anthropic details safeguards for Claude ahead of U.S. elections

Anthropic outlined measures intended to prevent election-related misuse of Claude, including restrictions on campaigning, lobbying and misinformation; automated enforcement backed by human review; targeted red-teaming; and large-scale evaluations. The company also directs election-related queries to current voting information and identifies Claude’s knowledge cutoff in its system prompt.

2 min read