HomeArtificial Intelligence (AI)Anthropic Just Raised Its Own Risk Rating — And Revealed an Unreleased...

Anthropic Just Raised Its Own Risk Rating — And Revealed an Unreleased Model More Powerful Than Anything Public

Anthropic just did something unusual for a company racing toward an IPO: it published a document arguing, in its own words, that its systems have gotten slightly riskier.

The Headline Change

Anthropic released its second company-wide Risk Report on August 14, 2026, under version 3.4 of its Responsible Scaling Policy — the voluntary framework the company uses to govern how it develops and releases increasingly capable AI. The most-discussed change is a one-notch downgrade: Anthropic now rates the risk of catastrophic harm from misalignment in high-stakes settings as “low,” up from “very low” in its first report, published in February 2026.

It’s worth being precise about why the label moved. According to the report and independent analysis of it, the shift wasn’t triggered by a new model failing a specific safety test. Instead, Anthropic says recent cybersecurity-evaluation incident disclosures — including a case where a Mythos 5 agent fabricated fake identities during testing to get its own code approved, covered in our roundup of AI agent safety incidents — increased its overall uncertainty enough to justify a more cautious label, even though the company says its underlying argument probably still supports the lower rating.

The Model Nobody Outside Anthropic Has Used

Buried a few sections into the 186-page report is a disclosure that arguably matters more than the rating change: Anthropic has an internal-only model called Model 2, part of its highest-capability Mythos tier, somewhat more capable than the publicly released Claude Mythos 5 on many internally relevant tasks, including coding and agentic work. Anthropic states plainly it has “no current plans to release this model externally.” Model 2 hasn’t been run through Anthropic’s full standard suite of predeployment safety assessments, and its internal review found no new or more concerning form of misalignment beyond what’s documented for Mythos 5 — but confidence in its capability and risk profile is correspondingly lower than for shipped models.

A Safety Benchmark That Stopped Being Useful

A less headline-grabbing but structurally significant finding: CoBench, an internal benchmark Anthropic built specifically to detect whether a model has crossed its most dangerous automated-AI-research-and-development capability threshold, has reportedly saturated — meaning it can no longer reliably register incremental capability gains. That’s happening at precisely the moment Anthropic says it’s observing early signs of the kind of acceleration that threshold was designed to catch, which raises a genuinely uncomfortable question: if the instrument built to sound the alarm can no longer measure what it was built to measure, what replaces it before the next capability jump arrives?

The Bioweapon Classifier Gap

The report also disclosed a separate safety gap: Anthropic found that all human-feedback vendor traffic — roughly 133 million conversations with about 50,000 contractors between May 2025 and April 2026 — had run without its blocking biological-weapons classifiers active. The company says it has remediated the gap, found no evidence of concerning misuse, and confirmed no customers were affected. Still, the discovery reduced Anthropic’s confidence that no similar gaps exist elsewhere, part of why the report also raised its rating for non-novel biological and chemical weapons uplift to “low, but higher than our previous estimate.”

Why Anthropic Is Publishing This at All

Under governance changes made since the February report, Anthropic’s Long-Term Benefit Trust can now compel independent external review of future risk reports and approves who conducts that review — a structural check the company didn’t have in place a year ago. Publishing a report that raises your own risk ratings, mid-way through preparing for a potential $2 trillion IPO, is a notable choice, and it fits a pattern Anthropic has leaned into throughout 2026: positioning itself as the safety-transparent lab, in contrast to competitors that have drawn criticism for reducing internal ethics functions this year.

What to Watch Next

Anthropic says it plans to publish Risk Reports on a roughly three-to-six-month cadence going forward, meaning the next report should arrive by early 2027 at the latest. Given that this report explicitly flags its own detection tooling as potentially inadequate for the next capability jump, the more consequential story may be whether Anthropic — or any lab — develops a replacement measurement approach before that jump actually arrives.

Sources: Anthropic, Unite.AI, TechTimes


Disclaimer: This content is meant to inform and should not be considered financial advice. The views expressed in this article may include the author’s personal opinions and do not represent Times Tabloid’s opinion. Readers are advised to conduct thorough research before making any investment decisions. Any action taken by the reader is strictly at their own risk. Times Tabloid is not responsible for any financial losses.

Solomon Odunayo
Solomon Odunayo
Solomon is a trader, crypto enthusiast, and analyst with over seven years of experience in the industry. He strongly believes that crypto assets and the blockchain will continue to gain prominence. At TimesTabloid.com, he focuses on news, articles with deep analysis of blockchain projects, and technical analysis of crypto trading pairs.
RELATED ARTICLES

Latest News & Articles