What Did Anthropic Propose?
Anthropic proposed an industry-wide framework for scoring AI jailbreak severity in July 2026, developed in partnership with Amazon, Microsoft, Google, and other enterprise partners under the Glasswing initiative. The framework defines a standardized taxonomy for categorizing AI model vulnerabilities, from minor output quality issues to critical safety failures.
This is the first major cross-vendor effort to establish a shared vocabulary for AI safety risk in enterprise deployments. It has direct implications for how B2B buyers evaluate AI vendors, how cybersecurity teams audit AI deployments, and how regulators may approach AI safety requirements in the near future.
Why Does Standardized Jailbreak Scoring Matter for Enterprise Buyers?
Until now, every AI vendor has defined safety and reliability in their own terms. Anthropic references constitutional AI. OpenAI cites alignment research. Google uses responsible AI principles. None of these are directly comparable to each other, which makes it hard for procurement teams to evaluate AI vendors on safety grounds.
A standardized jailbreak severity framework changes that. If adopted broadly, procurement teams at regulated enterprises, government agencies, and publicly traded companies will be able to evaluate AI vendors against a common safety benchmark, the same way they evaluate cybersecurity vendors against NIST or ISO frameworks.
What Does the Four-Level Severity Taxonomy Cover?
The Glasswing framework proposed by Anthropic scores AI vulnerabilities on a four-level scale:
Level 1 — Output quality failures: Hallucinations, inconsistent responses, or inaccurate information that degrades but does not endanger the user experience.
Level 2 — Policy compliance failures: Outputs that violate the AI vendor''s usage policies, such as generating restricted content categories.
Level 3 — Safety failures: Outputs that could cause harm if acted upon, such as incorrect medical or legal information presented with false confidence.
Level 4 — Critical safety failures: Outputs that bypass ethical constraints in ways that enable real-world harm, including providing detailed instructions for illegal activities or bypassing safety controls to access restricted capabilities.
For CISOs evaluating AI vendors, this taxonomy provides a structured way to ask: how many Level 3 and Level 4 incidents has this vendor documented in the past 12 months, and how were they remediated?
What Are the Three Immediate Implications for Enterprise Buyers?
AI safety benchmarking is becoming mandatory in enterprise procurement. RFPs for AI tooling are already including AI safety questions. A standardized severity framework accelerates this trend. Vendors who cannot demonstrate conformance with the Glasswing framework will face increased procurement friction in Q3 and Q4 2026.
The founding partners have a first-mover trust advantage. Amazon, Microsoft, Google, and Anthropic are all founding members. Enterprise buyers who standardize on these vendors for AI infrastructure are buying into a certified safety ecosystem, not just a product.
Niche AI vendors must respond or lose deals. Any AI vendor not participating in the severity framework will face procurement teams asking why they are not adopting the industry standard. This will start appearing in 2026 RFPs.
What Is the Cybersecurity Vendor Opportunity?
The jailbreak severity framework is good news for cybersecurity vendors in AI security and governance. As enterprises adopt the framework, they will need:
- Automated monitoring tools that detect and score jailbreak attempts against the Glasswing taxonomy in real time
- Audit trails documenting AI model behavior against the four-level severity framework
- Security operations workflows that escalate Level 3 and Level 4 events to the CISO and legal teams
This is a new product category conversation that can start with any enterprise account currently using or evaluating AI models in production workflows.
How Do You Get in Front of CISOs on This Conversation?
The fastest way to get CISOs talking about AI safety and jailbreak risk is to host an event that frames the conversation before the buyers know what they want. LinkedOtter''s event-led model works by identifying what buyers care about right now, AI safety governance being squarely in that category, and building a live roundtable that brings those buyers into a room.
From 1,266 prospects, clients consistently get 38 or more C-level executives to attend. The follow-up cadence targets the warmest attendees, generating 43 qualified meetings in 60 days from a single event cycle. See how it works and our events program for the structure.
Key Stats
- Anthropic proposed the jailbreak severity framework with Amazon, Microsoft, Google in July 2026
- The framework defines a four-level severity taxonomy from output quality failures to critical safety failures
- Glasswing, Anthropic''s critical infrastructure AI initiative, now covers 150+ organizations
- AI safety RFP questions are expected to accelerate in Q3 and Q4 2026 procurement cycles