A public commitment by an AI lab to only develop and deploy more capable models once specific safety evaluations and safeguards are in place.
Anthropic was among the first labs to publish one of these, tying its own model releases to concrete, pre-agreed safety thresholds rather than capability alone.
One of 60 free AI glossary terms
Plain-language definitions for the AI jargon you'll actually run into — no email needed, ever, for this section.
Browse the Full Glossary →