HomeArtificial Intelligence (AI)OpenAI's Smartest Model Just Got 14 Times Faster — Without Getting Any...

OpenAI’s Smartest Model Just Got 14 Times Faster — Without Getting Any Dumber

For most of the AI industry’s history, speed and intelligence have been a trade-off: want a fast response, drop down to a smaller, dumber model. OpenAI’s newest release argues that trade-off is now optional.

What Shipped

On August 13, OpenAI opened a limited preview of Ultrafast, a new service tier for its flagship model GPT-5.6 Sol, powered by chipmaker Cerebras. Ultrafast runs the model at up to 750 output tokens per second — as much as 14 times faster than GPT-5.6 Sol’s Standard processing tier — while OpenAI says it preserves the exact same intelligence and quality as the standard version. It’s available first through the OpenAI API to a select group of customers, with wider access planned as capacity grows.

The speed gain comes from Cerebras’s wafer-scale chip architecture, which eliminates a specific bottleneck — memory bandwidth between separate chip components — that has historically limited how fast large language models can generate text, regardless of how much raw computing power is thrown at the problem.

The Benchmark That Makes the Case

Cerebras and OpenAI published a striking demonstration: on Humanity’s Last Exam, a notoriously difficult 2,500-question benchmark spanning graduate-level chemistry, economics, and literature, GPT-5.6 Sol running on Ultrafast completed the entire question set in 11 hours and 11 minutes. According to Cerebras’s figures, that compares to roughly 78 hours for Anthropic’s Claude Fable 5 running at standard speed on the same benchmark — a gap of more than 60 hours on identical questions. Based on output speeds reported by Artificial Analysis, Cerebras also claims Ultrafast runs roughly 5 times faster than Claude Opus 4.8 in its Fast mode and 11 times faster than Claude Fable 5.

Those comparisons come from Cerebras and OpenAI themselves, worth treating as a starting point rather than independently verified fact — but the underlying claim, that wafer-scale chip architecture meaningfully outpaces standard GPU-based inference on raw token throughput, is consistent with Cerebras’s established technical reputation in the industry.

Why Speed, Not Just Intelligence, Is the New Battleground

OpenAI is framing Ultrafast around a specific insight: until now, anyone who needed genuinely real-time AI responses had to accept a smaller, less capable model in exchange for that speed. Ultrafast is pitched as dissolving that trade-off — delivering frontier-level reasoning at a speed previously reserved for much smaller, specialized models. OpenAI’s own internal teams have been testing the mode for incident response, where engineers can have logs, code changes, and reports analyzed by AI while an outage is still actively unfolding, rather than after the fact.

Early preview customers, including trading firm Jane Street, have echoed that framing — the speed jump doesn’t just make existing workflows marginally faster, it enables entirely different ways of working alongside the model in real time. That distinction matters most for agentic AI applications, where a model has to plan, act, and react across many sequential steps; latency at each step compounds quickly across a long agent run.

A Strategic Hedge Away From Nvidia

Beyond the product story, Ultrafast represents something else for OpenAI: a meaningful step away from total dependence on Nvidia hardware, building on the company’s existing ten-billion-dollar partnership with Cerebras and its parallel work developing custom silicon. As compute costs and chip supply constraints remain a defining pressure across the entire industry — see our coverage of the ongoing chip demand story — diversifying which hardware vendors can serve frontier models at scale has become a strategic priority in its own right, not just a performance optimization.

What to Watch Next

OpenAI hasn’t published pricing for Ultrafast, and running a flagship model at 14 times normal speed on specialized wafer-scale hardware is unlikely to come cheap — expect it to launch as a premium tier for latency-critical use cases rather than becoming the default way most people access GPT-5.6 Sol. As access expands beyond the current limited preview, the real test will be whether real-world agentic products built on Ultrafast show the kind of qualitative improvement OpenAI is promising, rather than just faster benchmark completion times.

Sources: OpenAI, HPCwire, The Next Web


Disclaimer: This content is meant to inform and should not be considered financial advice. The views expressed in this article may include the author’s personal opinions and do not represent Times Tabloid’s opinion. Readers are advised to conduct thorough research before making any investment decisions. Any action taken by the reader is strictly at their own risk. Times Tabloid is not responsible for any financial losses.

Solomon Odunayo
Solomon Odunayo
Solomon is a trader, crypto enthusiast, and analyst with over seven years of experience in the industry. He strongly believes that crypto assets and the blockchain will continue to gain prominence. At TimesTabloid.com, he focuses on news, articles with deep analysis of blockchain projects, and technical analysis of crypto trading pairs.
RELATED ARTICLES

Latest News & Articles