OpenAI Launches 'Ultrafast' Mode to Eliminate LLM Latency Bottlenecks

AI-generated image · Bay Street Wire
Powered by Cerebras, the new service tier accelerates GPT-5.6 Sol to 750 tokens per second, targeting mission-critical enterprise workflows.
OpenAI has introduced "Ultrafast," a new mode designed to accelerate its most powerful model, GPT-5.6 Sol, launching first via the OpenAI API. As reported by TechCrunch, this new tier can deliver up to 750 output tokens per second, operating at 14x the speed of standard processing.
Until now, enterprise users typically had to choose between high-intelligence frontier models and smaller, faster specialized models. OpenAI and chipmaker Cerebras, which powers the mode via its Wafer-Scale Engine architecture, claim Ultrafast resolves this tradeoff by providing frontier intelligence without quality compromise.
Cerebras reports that in head-to-head benchmarking on "Humanity's Last Exam" (HLE), GPT-5.6 Sol Ultrafast completed 2,500 PhD-level questions in 11 hours and 11 minutes, while Anthropic's Claude Fable 5 required 78 hours and 27 minutes. Additionally, Cerebras noted a 5.6x end-to-end speedup on the GDP-Val benchmark for economically valuable knowledge work.
OpenAI suggests the mode is specifically suited for high-stakes corporate workflows, including: * Incident response and root-causing production outages * Financial market analysis and economic modeling * Customer service and support * E-commerce * Cybersecurity detection and response
OpenAI researcher Jeffrey Wang stated the speed increase prevents the need to context-switch while waiting for tasks to finish. The feature is currently in limited preview for a select group of customers, with OpenAI stating access will expand as capacity grows.

