OpenAI’s Jalapeño Chip Sets New Inference Benchmarks
Reported by OpenAI Blog: Jalapeño’s first results show industry-leading speed and efficiency in AI inference. Analysis and context written by TickrWire.
OpenAI’s Jalapeño custom inference chip delivers 1.5‑3.6× better power efficiency and lower latency across major language models, outperforming commercial competitors.

- OpenAI’s Jalapeño chip delivers 1.5‑3.6× better power efficiency and lower latency on major language models.
- The chip’s design integrates compute, memory, networking, and software to minimize data movement and communication delays.
- Jalapeño outperforms commercial accelerators across the full InferenceX benchmark, achieving Pareto‑optimal performance.
- Deployment is slated for year‑end, with Gen 2 and Gen 3 already in development.
- Early adoption will reduce inference costs and improve operating leverage for OpenAI.
OpenAI has unveiled Jalapeño, its first custom silicon designed specifically for serving large‑scale language models. The company announced the chip in August 2026 after a nine‑month design cycle that leveraged AI‑driven design tools to accelerate tape‑out and verification. Jalapeño is positioned as the first generation of a multi‑generation platform that will be rolled out across OpenAI’s production infrastructure by year‑end.
In performance tests, Jalapeño achieved 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end‑to‑end latency than the leading commercial systems it was compared against. For highly interactive workloads, the chip delivered 2.1 to 4.1 times higher performance. The tests covered GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, with the chip rated at 700 W but sustaining 550 W or less during the workloads. These figures place Jalapeño on the Pareto frontier of the InferenceX benchmark, which measures the full request‑serving pipeline.
The chip’s design reflects OpenAI’s full‑stack philosophy. Rather than building a standalone accelerator, the team integrated model architecture, serving software, memory, networking, and rack‑scale systems around real inference workloads. This holistic approach allowed the designers to reduce data movement and communication delays, keeping the key KV cache local and minimizing idle time during token generation. The result is a balanced accelerator that excels at both the compute‑heavy prefill phase and the memory‑bandwidth‑heavy decode phase.
Compared to existing hardware, which typically forces a trade‑off between throughput and latency, Jalapeño offers a unified solution that improves both metrics. In the InferenceX benchmark, the chip outperformed commercial accelerators across the entire operating range, from high‑throughput batch serving to low‑latency interactive use. The chip’s network architecture keeps the entire workload within a single connected system, further reducing latency and improving energy efficiency.
Despite the impressive gains, Jalapeño is still in the early stages of deployment. Production qualification, software maturation, and validation across a broader set of models remain on the roadmap. Each new model family requires custom kernels and optimizations, and while AI tools like Codex can accelerate this process, human oversight is still necessary. Manufacturing costs and integration complexity also pose challenges as the chip moves from prototype to mass production.
OpenAI plans to roll out Jalapeño across its compute infrastructure by the end of the year, with Gen 2 and Gen 3 already in development. The company will continue to use NVIDIA and other partners for training workloads while deploying Jalapeño for inference. The improved efficiency is expected to lower the cost of serving advanced models, increase operating leverage, and enable faster iteration on new use cases.
In summary, Jalapeño represents a significant step forward in AI inference hardware. By combining higher throughput, lower latency, and better power efficiency in a single architecture, it sets a new benchmark for custom silicon and demonstrates the value of a full‑stack design approach. The chip’s deployment will likely influence the broader industry’s direction toward integrated, AI‑optimized hardware solutions.
Enables faster, more efficient inference pipelines and showcases AI‑driven hardware design.
Reduces operational costs for large‑scale AI services and improves user experience.
Signals OpenAI’s continued investment in custom silicon, potentially increasing long‑term profitability.
Provides a real‑world case study of full‑stack AI system design.
Highlights progress toward more accessible and affordable AI applications.
- prefill
- The phase where the model processes the input prompt before generating tokens.
- decode
- The token‑by‑token generation phase during inference.
- KV cache
- Key‑value memory that stores intermediate states for efficient token generation.
- mixed TPS/kW
- Throughput measured as tokens per second per kilowatt, indicating energy efficiency.
HardwareOpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
OpenAI shared initial benchmark data for its Jalapeño inference chip at the Hot Chips conference, demonstrating higher token generation and energy efficiency than current market standards.
AI ToolsAccel-backed Keenable is indexing the web for AI agents
Keenable has emerged from stealth with $26 million in funding to provide a specialized web search index designed specifically for AI agents rather than human users.
AI ToolsNvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated
Nvidia has moved its Groq 3 LPX inference accelerator into full production, reporting 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras, but the comparison depends on accelerator count.
SecurityUkraine opens its massive labeled battlefield dataset to British firms in a landmark AI weapons partnership
Ukraine has granted the United Kingdom access to its Avengers Labs platform, providing foreign tech companies with millions of annotated combat images to train military artificial intelligence.
AI Tools‘The world seems to be ready’: An interview with OpenAI head of product Thibault Sottiaux
OpenAI’s head of product Thibault Sottiaux discusses the launch of ChatGPT Work, a platform for white-collar workers to use AI agents, its $20/month pricing, and the challenges of scaling AI adoption.
RoboticsI spent a day at a robot “carnival” in Shanghai. Here’s what I saw.
A recent robotics festival in Shanghai highlighted China's rapid commercial progress in humanoid systems, where local firms now dominate global delivery numbers.