← Back

safety

More details on Fable 5’s cyber safeguards and our jailbreak framework

Anthropic has re-deployed Claude Fable 5 with enhanced cybersecurity safeguards and introduced an early draft of a jailbreak severity framework to evaluate risks associated with model misuse.

AS1 NewsSource: anthropic.com

ai-safetymodel-safeguardsjailbreak-frameworkcybersecurityai-regulationanthropic
Anthropic$2,054.92-0.61%

Claude Fable 5 has been made available globally, accompanied by detailed information on its cybersecurity safeguards. These include safety classifiers designed to detect and block dangerous uses, particularly in cybersecurity contexts. The classifiers categorize requests into four levels of risk, from potentially benign to overtly harmful, allowing for nuanced control over model outputs. The safeguards are part of a broader safety system that also includes access controls, model training, and offline monitoring.

A key focus of the update is on AI jailbreaks—methods that prompt models to bypass safeguards. Anthropic has developed an early draft of a jailbreak severity framework in collaboration with its Glasswing partners. This framework aims to standardize how the severity of jailbreaks is assessed, facilitating clearer communication between AI developers, regulators, and other stakeholders about the risks posed by specific jailbreak techniques.

The company emphasizes that all cybersecurity-related AI capabilities are dual-use, meaning they can be employed for both defensive and malicious purposes. Consequently, the safeguards are designed to block high-risk activities such as defense evasion, data exfiltration, and automatic vulnerability discovery, especially when these actions could be exploited maliciously. The safeguards are adaptable, with the safety margin adjustable to balance between enabling beneficial uses and preventing harm.

Anthropic has also launched a HackerOne program to invite security researchers to report potential jailbreaks, fostering a collaborative approach to safety. While the safeguards aim to prevent misuse, the company recognizes the importance of allowing cybersecurity professionals to use the model for legitimate defense activities, such as vulnerability scanning and testing, under appropriate controls.

This development underscores ongoing efforts within the AI community to improve safety measures, particularly in sensitive areas like cybersecurity, where capabilities can be easily misused. The introduction of a jailbreak severity framework represents a step toward more transparent and standardized risk assessment practices, which could influence future safety protocols and regulatory discussions.

neutral

Enhances safety protocols for AI models, potentially influencing industry standards and regulatory policies.