Waiting for an AI to finish its thought can feel like watching a slow-motion typist. Now, OpenAI is trying to change that experience by launching a new high-speed version of its most powerful AI model, GPT-5.6 Sol.
OpenAI has introduced a new service mode called Ultrafast that runs its flagship model, GPT-5.6 Sol, at 14 times the speed of its standard version. This means the model can now generate up to 750 tokens per second. Tokens are essentially the small building blocks of text—like words or parts of words—that an AI produces as it answers your request. To achieve this, OpenAI partnered with Cerebras, a company that designs specialized hardware chips meant to handle the intense, rapid calculations that AI requires. This feature is currently in a limited testing phase and is aimed primarily at businesses using AI for high-speed tasks like financial data analysis or customer service.
The trade-off between speed and brainpower
To understand why this is a change, you have to look at the trade-off that has existed until now. Think of an AI model like a team of researchers. A powerful, intelligent model is like a team of experts that can solve complex, nuanced problems, but it takes them a long time to read, think, and write out their findings. A smaller, less capable model is like a speedy intern; they answer questions instantly, but they often lack the deep knowledge or reasoning skills needed for difficult tasks. In the past, companies usually asked users to pick: do you want the expert, or do you want the speed? OpenAI is attempting to bridge this gap by using highly specialized computer chips that can process these expert-level calculations much faster than the standard hardware usually used by these systems.
For most people, this is a sign that the friction of waiting for digital assistants is likely to disappear. In a professional setting, this could mean that automated customer service chats feel truly instantaneous, or that financial tools can scan and summarize data in a blink of an eye. If these high-speed tools perform reliably, we might stop thinking of AI as a tool that generates text at its own pace and start seeing it as a real-time partner in our daily work. The real test will be whether this speed comes with any drop in accuracy, as speed sometimes encourages shortcuts that can lead to errors.
Liked this one? The next lands at breakfast.
Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.
By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy