OpenAI unveils custom AI inference chip Jalapeño
Reported by OpenAI Blog: OpenAI says its Jalapeño chip can power faster AI responses than the competition. Analysis and context written by TickrWire.
Evolving story · 2 updatesOpenAI Custom Jalapeño Chip DevelopmentTimeline →OpenAI introduced its first custom inference chip, Jalapeño, which on the InferenceX benchmark delivered higher peak throughput per kilowatt and lower token latency than competing commercial systems, and performed well across several model families.

- OpenAI’s Jalapeño chip outperformed commercial accelerators on the InferenceX benchmark, delivering higher throughput per kilowatt and lower latency.
- The chip’s performance was validated across multiple model families, including GPT‑OSS 120B, DeepSeek R1, and Kimi K2.
- OpenAI’s integrated stack strategy combines custom silicon, software, and cloud partners to optimize cost and capability for diverse AI workloads.
OpenAI’s chief financial officer, Sarah Friar, used a recent blog post to explain the company’s holistic approach to scaling artificial intelligence. At the centre of that strategy is a new piece of hardware – the Jalapeño chip – which the firm says marks its first foray into custom inference silicon. The announcement highlighted measured results from a public benchmark called InferenceX, where Jalapeño ran a 120‑billion‑parameter GPT‑OSS model and achieved higher peak throughput per kilowatt and lower token latency than the commercial accelerators it was compared against. The chip also posted strong numbers on other model families such as DeepSeek R1 and Kimi K2, suggesting the gains are not limited to a single architecture.
The Jalapeño chip is presented as a lever for OpenAI to gain tighter control over both performance and cost. By co‑designing the model, serving software, chip, memory, and networking layers, the company claims it can improve throughput, latency, energy efficiency, and overall economics as a single system rather than as isolated components. Friar noted that this first‑generation silicon is already in production, and that future generations are in development, indicating a longer‑term roadmap for in‑house hardware.
OpenAI’s compute strategy, according to the post, rests on an integrated stack that spans data‑center facilities, custom chips, frontier models, developer platforms, consumer and enterprise products, and AI‑native devices. Each layer reinforces the others: better software extracts more performance from hardware, while more capable models enable richer products that generate additional usage signals, which in turn feed back into system improvements. The company emphasizes that different workloads – such as frontier training, high‑volume inference, and always‑on agents – have distinct requirements across chips, software, networks, power, and latency, and that its portfolio is designed to meet those varied demands.
OpenAI’s portfolio of partners includes long‑standing collaborators like Microsoft’s compute infrastructure and NVIDIA’s GPUs, as well as newer contributors such as AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank. By maintaining a mix of premium systems for capability‑heavy tasks and optimized setups for cost‑sensitive scaling, OpenAI aims to stay on what it calls the Pareto frontier – the optimal balance of capability, speed, reliability, efficiency, and cost for each workload. The post argues that preserving choice across providers and hardware types lets the company direct demand toward the most cost‑effective solutions while retaining pricing discipline as market conditions evolve.
Beyond the chip itself, OpenAI highlighted a data‑center project named Camellia in Georgia. The initiative is described as a purpose‑built facility that aligns its design with customer workloads, creates local jobs, supports nearby businesses, and incorporates a closed‑loop water system. An independent public audit is said to verify the project’s commitments each year, underscoring OpenAI’s focus on sustainable and accountable infrastructure.
The post concludes by linking hardware efficiency to broader economic effects. Using an internal metric called the Artificial Analysis Coding Agent Index, OpenAI reported that its GPT‑5.6 Sol model achieved a new high on reasoning tasks while using 54 % fewer output tokens than a leading competitor. The company frames this as evidence that more useful intelligence per dollar enables new use cases – from contract review to live financial scenario planning – and triggers a Jevons‑paradox effect where greater efficiency expands overall consumption. OpenAI asserts that the compounding advantage of better technology, lower costs, and reinvested growth fuels a virtuous cycle of continued progress in AI capabilities and accessibility.
Custom silicon can lower inference costs and latency for applications built on OpenAI models.
More efficient inference translates to cheaper AI services and the ability to scale new use cases.
A successful first‑party chip signals OpenAI’s move toward greater vertical integration and potential cost advantages.
Provides a concrete example of how hardware and software co‑design drives AI performance improvements.
- InferenceX
- A public benchmark that measures AI inference performance, often using large language models.
- Jalapeño
- OpenAI’s first custom inference chip designed to improve throughput and energy efficiency.
- Jevons paradox
- An economic observation that increased efficiency can lead to higher overall consumption of a resource.
- Pareto frontier
- The set of optimal trade‑offs where improving one metric would worsen another, used here to balance capability, speed, and cost.
AI bias estimate: The source is an OpenAI self‑authored post, so it emphasizes positive outcomes and may understate limitations or competitive challenges. (Automated estimate, not a definitive judgement.)
- OpenAI says its Jalapeño chip can power faster AI responses than the competition ↗
- OpenAI details Jalapeño AI chip, with 700W TDP - Data Center Dynamics ↗
- OpenAI Jalapeño: Better Than Nvidia Blackwell - SemiAnalysis ↗
- OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks ↗
- The full stack behind abundant intelligence ↗
- OpenAI Says Its Jalapeño AI Chip Is Better Than Nvidia’s Blackwell - The Information ↗
- OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show ↗
- OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show - TechCrunch ↗
HardwareMeta AI Introduces MetaRoCE: A Clean-Sheet RDMA Transport Built for AI-Scale Ethernet
HardwareOpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
HardwareApple Debuts M6 and M5 Ultra Chips for a Big Leap in AI Compute
HardwareCerebras unveils CS-4 with double the performance on the same chip
HardwareRayNeo's new AI glasses skip the camera, focus on text overlays
FundingIndia’s Ringg gets backing from Peak XV as it pushes voice AI past the phone call
Indian voice AI startup Ringg has raised $10 million in a Series A extension led by Peak XV Partners, bringing its total funding in the round to $15.5 million.
RoboticsRobotics startup Generalist reaches $3B valuation, sources say
Robotics startup Generalist secured a nearly $200 million funding extension led by 8VC, lifting its valuation to $3 billion just months after a major Series B round.
BusinessOpenAI loses a top data center exec as stream of high-profile departures continues
OpenAI’s head of data centers, Chris Malone, has left the company as part of a broader executive exodus, raising questions about leadership stability ahead of a planned IPO.
AI ResearchAI Method Reveals What Genomic Models Learn From DNA and Exposes Hidden Experimental Bias
Researchers at the Stowers Institute introduced PISA, a pairwise influence by sequence attribution method that visualizes, at single‑base resolution, what deep‑learning models learn from DNA and can strip experimental bias from MNase‑seq data.
AI ToolsPerplexity Ships Portable Computer on NVIDIA DGX Spark: Local Harness, OS-Enforced Sandbox, and Zero Per-Token Cost for Local Steps
Perplexity introduced Portable Computer, a bundled local‑first AI agent system that runs on NVIDIA DGX Spark and eliminates per‑token fees for on‑device processing.
FundingStability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding
Stability AI announced a $76 million Series B round, bringing its total funding to $232 million. Investors include Universal Music Group, Sony Music, Warner Music, Electronic Arts, AMD Ventures and Pacific Alliance Ventures.