GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras
OpenAI introduces Ultrafast mode for GPT-5.6 Sol, delivering up to 750 tokens per second using Cerebras hardware, tripling its inference speed.
- Ultrafast mode delivers 750 tokens per second for GPT-5.6 Sol, a 14x speedup over baseline performance.
- The mode is powered by Cerebras’ wafer-scale hardware, part of a $10 billion OpenAI partnership.
- Ultrafast joins Standard and Fast tiers, creating a three-tier pricing structure focused on inference speed.
- This move could intensify competition among AI providers to optimize hardware for high-throughput inference.
OpenAI has launched Ultrafast mode for its GPT-5.6 Sol model, a significant performance upgrade that pushes inference speeds to 750 tokens per second. This is powered by Cerebras Systems’ wafer-scale hardware, part of their $10 billion partnership with OpenAI. The new mode joins existing Standard and Fast tiers, forming a three-tier pricing structure where speed itself becomes a product feature.
The Ultrafast mode represents a 14x improvement over baseline GPT-5.6 Sol performance, addressing a critical bottleneck for real-time AI applications. Cerebras’ custom silicon is designed to handle massive parallel processing, making it ideal for high-throughput inference workloads. This move aligns with OpenAI’s broader strategy to monetize inference speed as a premium service, particularly for enterprise and latency-sensitive use cases.
Industry observers note that this could pressure competitors like Anthropic and Mistral to accelerate their own hardware partnerships. The pricing tiers suggest OpenAI is segmenting the market by performance, potentially increasing revenue per query while catering to diverse customer needs.
Developers can now build real-time AI applications with significantly reduced latency using the Ultrafast tier.
Enterprises can choose pricing tiers based on performance needs, optimizing cost for latency-sensitive workloads.
Demonstrates OpenAI’s commitment to monetizing performance differentiation, potentially boosting revenue per inference.
Accelerates the shift toward real-time AI interactions, making advanced models more practical for everyday use.
- tokens per second
- A metric measuring how many tokens (words or word fragments) an AI model can generate in one second.
- wafer-scale hardware
- Custom silicon chips designed to process entire AI workloads on a single wafer, reducing latency and increasing throughput.
AI eyes in the sky: New satellites and artificial intelligence are transforming wildfire detection - AccuWeather
HardwareFirst test flight of largest all-electric aircraft used just $5 of electricity
HardwareOrganic-looking brake assemblies debut on new Czinger 21C Spyder
HardwareOpenAI and Cerebras Bring GPT-5.6 Sol Ultrafast to Enterprise Inference
HardwareWe've flown a radiation-blocking vest to the Moon and back, and it worked

The "AI" Badge Doesn't Measure What You Think It Does
Anthropic joined the EU AI Act's transparency code, but the 'AI' badge may mislead users about content authenticity.
AI could help fossil fuel companies create more emissions - grist.org
A new report suggests AI tools could enable fossil fuel firms to extract and burn more hydrocarbons, worsening climate impact.
AI ResearchHow I Built a Real-Time Multilingual AI Voice Tutor for Bharat (And Solved the 55ms Latency Problem)
A developer built a real-time multilingual AI voice tutor, solving a 55ms latency issue. The tutor is designed for Bharat, indicating potential for broader language support.
HBCU Love: NCCU opens nation’s first HBCU artificial intelligence institute building - Texas Metro News
North Carolina Central University has opened the first artificial intelligence institute building dedicated to HBCUs, marking a historic milestone for historically Black colleges.
BusinessAI-generated books are flooding Amazon and tanking sales for human authors
A new study reveals AI-generated titles now make up 20 percent of Amazon's self-published catalog. This influx correlates with declining revenue for human authors across most genres.
Artificial Intelligence in Predicting Systemic Complications From Retinal Findings: A New Frontier in Precision Medicine - Cureus
A new AI model analyzes retinal images to forecast serious systemic complications, marking a leap in precision medicine.