HardwareAug 14, 2026, 2:21 PM

GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras

30-second summary

OpenAI introduces Ultrafast mode for GPT-5.6 Sol, delivering up to 750 tokens per second using Cerebras hardware, tripling its inference speed.

TickrWire
GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras
Key takeaways
  • Ultrafast mode delivers 750 tokens per second for GPT-5.6 Sol, a 14x speedup over baseline performance.
  • The mode is powered by Cerebras’ wafer-scale hardware, part of a $10 billion OpenAI partnership.
  • Ultrafast joins Standard and Fast tiers, creating a three-tier pricing structure focused on inference speed.
  • This move could intensify competition among AI providers to optimize hardware for high-throughput inference.
Full story

OpenAI has launched Ultrafast mode for its GPT-5.6 Sol model, a significant performance upgrade that pushes inference speeds to 750 tokens per second. This is powered by Cerebras Systems’ wafer-scale hardware, part of their $10 billion partnership with OpenAI. The new mode joins existing Standard and Fast tiers, forming a three-tier pricing structure where speed itself becomes a product feature.

The Ultrafast mode represents a 14x improvement over baseline GPT-5.6 Sol performance, addressing a critical bottleneck for real-time AI applications. Cerebras’ custom silicon is designed to handle massive parallel processing, making it ideal for high-throughput inference workloads. This move aligns with OpenAI’s broader strategy to monetize inference speed as a premium service, particularly for enterprise and latency-sensitive use cases.

Industry observers note that this could pressure competitors like Anthropic and Mistral to accelerate their own hardware partnerships. The pricing tiers suggest OpenAI is segmenting the market by performance, potentially increasing revenue per query while catering to diverse customer needs.

Sponsored
Why this matters
Developers

Developers can now build real-time AI applications with significantly reduced latency using the Ultrafast tier.

Businesses

Enterprises can choose pricing tiers based on performance needs, optimizing cost for latency-sensitive workloads.

Investors

Demonstrates OpenAI’s commitment to monetizing performance differentiation, potentially boosting revenue per inference.

Everyone

Accelerates the shift toward real-time AI interactions, making advanced models more practical for everyday use.

Glossary
tokens per second
A metric measuring how many tokens (words or word fragments) an AI model can generate in one second.
wafer-scale hardware
Custom silicon chips designed to process entire AI workloads on a single wafer, reducing latency and increasing throughput.
Sources · 1
Read next
More stories
TickrWireAI News Intelligence

We aggregate, verify, summarise and explain the latest artificial intelligence news from open, legal sources.

Daily AI digest

Top AI stories, summarised, in your inbox each morning.

© 2026 TickrWire. Summaries and analysis are AI-generated and may contain errors.