OpenAI Details Jalapeño Chip Performance
Reported by TechCrunch AI: OpenAI says its Jalapeño chip can power faster AI responses than the competition. Analysis and context written by TickrWire.
OpenAI shared initial benchmark data for its Jalapeño inference chip at the Hot Chips conference, demonstrating higher token generation and energy efficiency than current market standards.

- OpenAI presented initial benchmark results for its custom Jalapeño chip at the Hot Chips conference.
- The chip outperformed Nvidia Blackwell systems in token generation and throughput per kilowatt on the InferenceX benchmark.
- OpenAI plans initial small volume deployments by late 2026, followed by broader scaling in 2027.
- The chip was developed in collaboration with Broadcom and aims to minimize bottlenecks during the prefill and communication phases.
OpenAI has provided a more detailed overview of its custom silicon project known as Jalapeño, sharing the first set of benchmark evaluations during a presentation at the Hot Chips conference. According to the data shared by the organization, the custom processor underwent testing using the InferenceX benchmark suite developed by Semianalysis. The results indicate that the hardware achieved higher token generation counts per user alongside greater processing throughput per kilowatt when measured against current top tier inference solutions available on the market.
Richard Ho, who leads the hardware division at OpenAI, emphasized the scale of the performance leap during a press briefing. He noted that the chip is capable of handling heavier AI workloads while consuming less electrical power, while simultaneously decreasing response times for end users. The objective behind the design is to establish a balance between high capacity service delivery and minimal latency for high volume applications.
The benchmark comparisons explicitly measured the Jalapeño system against Nvidia Blackwell hardware configurations. However, industry analysts note that competitive hardware will likely continue to evolve by the time OpenAI reaches full production status. During the briefing, Ho provided a projected timeline estimating that initial low volume deployments of the chip will begin near the end of 2026, with broader scaling anticipated throughout 2027.
Development of the Jalapeño architecture began with an initial public disclosure last October, driven by a partnership between OpenAI and Broadcom. Internal tools and AI models developed by OpenAI were also utilized to assist in designing the silicon. The company views Jalapeño as the foundation for a multigenerational hardware strategy, aiming to co-design future software models, AI applications, memory systems, and silicon processors simultaneously.
This vertical integration strategy allowed the engineering team to target specific operational bottlenecks that typically degrade performance during the inference cycle. Specifically, the architecture focuses on reducing friction points that occur during the prefill and communication stages of token generation. OpenAI stated that traditional hardware often struggles with data transfer overhead during these particular phases.
To overcome these hardware limitations, the Jalapeño design emphasizes localized data movement and storage optimization. The system architecture ensures that critical model parameters, including the key value cache required during text generation, remain localized while dynamic compute, memory, and networking resources are allocated for each distinct phase of processing.
Despite the promising early metrics, several commercial and logistical questions remain regarding the eventual rollout. Scaling custom silicon manufacturing requires navigating complex supply chains and substantial capital expenditures. Furthermore, the competitive landscape will not stand still, as existing hardware vendors continuously update their product lines to improve energy efficiency and raw compute power.
Market observers will be monitoring how OpenAI manages the transition from external hardware dependencies to proprietary silicon infrastructure. The success of the Jalapeño platform will depend heavily on whether the actual deployment schedule holds and whether the performance advantages observed in controlled benchmark tests translate effectively into large scale production environments.
Custom hardware optimization could eventually alter how large models are served and queried.
Lower inference costs per kilowatt may improve the unit economics of running large scale AI services.
OpenAI is executing a vertical integration strategy to reduce reliance on third party silicon vendors.
Custom chip development by major labs helps address future energy and infrastructure constraints.
- InferenceX
- A benchmark suite created by Semianalysis used to measure the performance of AI inference hardware.
- KV cache
- Key-value cache memory stored during generative AI processing to avoid recalculating previous tokens.
AI bias estimate: The report relies primarily on metrics provided directly by OpenAI and Semianalysis without independent third party verification. (Automated estimate, not a definitive judgement.)
- OpenAI says its Jalapeño chip can power faster AI responses than the competition ↗
- OpenAI details Jalapeño AI chip, with 700W TDP - Data Center Dynamics ↗
- OpenAI’ Jalapeño: Better Than Nvidia Blackwell - SemiAnalysis ↗
- OpenAI’s Jalapeño AI Chip Outperforms Nvidia Blackwell in Early Tests - citybiz ↗
- OpenAI says its Jalapeño chip bests Nvidia, others - Axios ↗
- OpenAI says its Jalapeño chip can power faster AI responses than the competition - The Verge ↗
- OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show ↗
- OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show - TechCrunch ↗
- OpenAI releases first performance data for Jalapeño AI chip - Investing.com ↗
HardwareCerebras unveils CS-4 with double the performance on the same chip
HardwareRayNeo's new AI glasses skip the camera, focus on text overlays
HardwareTrump's space transportation policy calls for new spaceport on federal land
HardwareMotorola's GrapheneOS phones will launch in 2027 priced higher than Pixels
HardwareData center opposition surged from 42 to 75 percent in just one year, survey finds
AI ToolsAccel-backed Keenable is indexing the web for AI agents
Keenable has emerged from stealth with $26 million in funding to provide a specialized web search index designed specifically for AI agents rather than human users.
AI ToolsNvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated
Nvidia has moved its Groq 3 LPX inference accelerator into full production, reporting 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras, but the comparison depends on accelerator count.
SecurityUkraine opens its massive labeled battlefield dataset to British firms in a landmark AI weapons partnership
Ukraine has granted the United Kingdom access to its Avengers Labs platform, providing foreign tech companies with millions of annotated combat images to train military artificial intelligence.
AI Tools‘The world seems to be ready’: An interview with OpenAI head of product Thibault Sottiaux
OpenAI’s head of product Thibault Sottiaux discusses the launch of ChatGPT Work, a platform for white-collar workers to use AI agents, its $20/month pricing, and the challenges of scaling AI adoption.
RoboticsI spent a day at a robot “carnival” in Shanghai. Here’s what I saw.
A recent robotics festival in Shanghai highlighted China's rapid commercial progress in humanoid systems, where local firms now dominate global delivery numbers.
SecurityTaiwanese cybersecurity firm warns that AI tools have more than doubled Chinese state-backed cyberattacks
Taiwanese cybersecurity firm TeamT5 reports that Chinese state-backed hacking groups have more than doubled their attack frequency after adopting AI models such as DeepSeek.