← AI News

August 14, 2026 · OpenAI

OpenAI Previews Ultrafast: GPT-5.6 Sol at Up to 750 Tokens Per Second with Cerebras

My take: Speed has always been the constraint that separated the most intelligent models from the ones you could actually use in real time. OpenAI just introduced Ultrafast, a new API service tier that runs GPT-5.6 Sol on Cerebras hardware at up to 750 tokens per second, up to 14 times faster than standard processing. In practice, that collapses wait times from several seconds to fractions of a second.

For anyone building products where the model has to respond while the user is still waiting, this changes what is possible: live customer support, voice applications, real-time financial analysis, or any workflow that cannot afford latency. This is not a personal use improvement; it is what makes the most intelligent model also the fastest one.

Access starts as a limited preview for a select group of API customers, expanding gradually. If you are building something where model latency still blocks the user experience, this is the moment to apply for access.

Read at the source: OpenAI ↗

Want to use these tools? See the unbiased reviews or back to the news.