HardwareAug 25, 2026, 2:22 PM

OpenAI Details Jalapeño Chip Performance

TickrWire Editorial Desk·Aug 25, 2026, 2:22 PM·2 min read AI-assisted, human-reviewed

Reported by TechCrunch AI: OpenAI says its Jalapeño chip can power faster AI responses than the competition. Analysis and context written by TickrWire.

30-second summary

OpenAI shared initial benchmark data for its Jalapeño inference chip at the Hot Chips conference, demonstrating higher token generation and energy efficiency than current market standards.

TickrWire
OpenAI Details Jalapeño Chip Performance
Key takeaways
  • OpenAI presented initial benchmark results for its custom Jalapeño chip at the Hot Chips conference.
  • The chip outperformed Nvidia Blackwell systems in token generation and throughput per kilowatt on the InferenceX benchmark.
  • OpenAI plans initial small volume deployments by late 2026, followed by broader scaling in 2027.
  • The chip was developed in collaboration with Broadcom and aims to minimize bottlenecks during the prefill and communication phases.
Full story

OpenAI has provided a more detailed overview of its custom silicon project known as Jalapeño, sharing the first set of benchmark evaluations during a presentation at the Hot Chips conference. According to the data shared by the organization, the custom processor underwent testing using the InferenceX benchmark suite developed by Semianalysis. The results indicate that the hardware achieved higher token generation counts per user alongside greater processing throughput per kilowatt when measured against current top tier inference solutions available on the market.

Richard Ho, who leads the hardware division at OpenAI, emphasized the scale of the performance leap during a press briefing. He noted that the chip is capable of handling heavier AI workloads while consuming less electrical power, while simultaneously decreasing response times for end users. The objective behind the design is to establish a balance between high capacity service delivery and minimal latency for high volume applications.

The benchmark comparisons explicitly measured the Jalapeño system against Nvidia Blackwell hardware configurations. However, industry analysts note that competitive hardware will likely continue to evolve by the time OpenAI reaches full production status. During the briefing, Ho provided a projected timeline estimating that initial low volume deployments of the chip will begin near the end of 2026, with broader scaling anticipated throughout 2027.

Development of the Jalapeño architecture began with an initial public disclosure last October, driven by a partnership between OpenAI and Broadcom. Internal tools and AI models developed by OpenAI were also utilized to assist in designing the silicon. The company views Jalapeño as the foundation for a multigenerational hardware strategy, aiming to co-design future software models, AI applications, memory systems, and silicon processors simultaneously.

This vertical integration strategy allowed the engineering team to target specific operational bottlenecks that typically degrade performance during the inference cycle. Specifically, the architecture focuses on reducing friction points that occur during the prefill and communication stages of token generation. OpenAI stated that traditional hardware often struggles with data transfer overhead during these particular phases.

To overcome these hardware limitations, the Jalapeño design emphasizes localized data movement and storage optimization. The system architecture ensures that critical model parameters, including the key value cache required during text generation, remain localized while dynamic compute, memory, and networking resources are allocated for each distinct phase of processing.

Despite the promising early metrics, several commercial and logistical questions remain regarding the eventual rollout. Scaling custom silicon manufacturing requires navigating complex supply chains and substantial capital expenditures. Furthermore, the competitive landscape will not stand still, as existing hardware vendors continuously update their product lines to improve energy efficiency and raw compute power.

Market observers will be monitoring how OpenAI manages the transition from external hardware dependencies to proprietary silicon infrastructure. The success of the Jalapeño platform will depend heavily on whether the actual deployment schedule holds and whether the performance advantages observed in controlled benchmark tests translate effectively into large scale production environments.

Why this matters
Developers

Custom hardware optimization could eventually alter how large models are served and queried.

Businesses

Lower inference costs per kilowatt may improve the unit economics of running large scale AI services.

Investors

OpenAI is executing a vertical integration strategy to reduce reliance on third party silicon vendors.

Everyone

Custom chip development by major labs helps address future energy and infrastructure constraints.

Glossary
InferenceX
A benchmark suite created by Semianalysis used to measure the performance of AI inference hardware.
KV cache
Key-value cache memory stored during generative AI processing to avoid recalculating previous tokens.

AI bias estimate: The report relies primarily on metrics provided directly by OpenAI and Semianalysis without independent third party verification. (Automated estimate, not a definitive judgement.)

Sources · 9
Read next
More stories
Accel-backed Keenable is indexing the web for AI agentsAI Tools

Accel-backed Keenable is indexing the web for AI agents

Keenable has emerged from stealth with $26 million in funding to provide a specialized web search index designed specifically for AI agents rather than human users.

Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicatedAI Tools

Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated

Nvidia has moved its Groq 3 LPX inference accelerator into full production, reporting 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras, but the comparison depends on accelerator count.

Ukraine opens its massive labeled battlefield dataset to British firms in a landmark AI weapons partnershipSecurity

Ukraine opens its massive labeled battlefield dataset to British firms in a landmark AI weapons partnership

Ukraine has granted the United Kingdom access to its Avengers Labs platform, providing foreign tech companies with millions of annotated combat images to train military artificial intelligence.

‘The world seems to be ready’: An interview with OpenAI head of product Thibault SottiauxAI Tools

‘The world seems to be ready’: An interview with OpenAI head of product Thibault Sottiaux

OpenAI’s head of product Thibault Sottiaux discusses the launch of ChatGPT Work, a platform for white-collar workers to use AI agents, its $20/month pricing, and the challenges of scaling AI adoption.

I spent a day at a robot “carnival” in Shanghai. Here’s what I saw.Robotics

I spent a day at a robot “carnival” in Shanghai. Here’s what I saw.

A recent robotics festival in Shanghai highlighted China's rapid commercial progress in humanoid systems, where local firms now dominate global delivery numbers.

Taiwanese cybersecurity firm warns that AI tools have more than doubled Chinese state-backed cyberattacksSecurity

Taiwanese cybersecurity firm warns that AI tools have more than doubled Chinese state-backed cyberattacks

Taiwanese cybersecurity firm TeamT5 reports that Chinese state-backed hacking groups have more than doubled their attack frequency after adopting AI models such as DeepSeek.