JetSpec: 9.64x LLM Speedup with Parallel Tree Drafting
Reported by the original publisher: [Research] JetSpec: Speculative Decoding with Parallel Tree Drafting Enables up to 9.64x Lossless LLM Inference Speedup with more than 1000TPS. Analysis and context written by TickrWire.
JetSpec introduces a novel speculative decoding method using parallel tree drafting to achieve up to 9.64x lossless speedup in LLM inference, reaching over 1000 TPS on a single B200 GPU.

- JetSpec achieves up to 9.64x lossless speedup in LLM inference using parallel tree drafting for speculative decoding.
- Performance gains demonstrated on MATH-500 (9.64x) and open-ended chat (4.58x) benchmarks.
- Throughput exceeds 1000 TPS on a single B200 GPU with CUDA optimizations.
- Method maintains lossless generation while optimizing drafting cost and quality.
- Builds on prior speculative decoding work but introduces parallel tree drafting for efficiency.
Researchers have developed JetSpec, a speculative decoding framework that optimizes both drafting cost and quality through causal parallel tree drafting. This method enables lossless inference speedups of up to 9.64x on MATH-500 and 4.58x on open-ended chat benchmarks. By leveraging CUDA graph and kernel optimizations, JetSpec achieves throughput exceeding 1000 tokens per second (TPS) on a single NVIDIA B200 GPU. The approach addresses prior limitations of speculative decoding by improving drafting efficiency without sacrificing accuracy.
Provides a practical, high-performance speculative decoding method for faster LLM inference without accuracy loss, with open-source potential.
Enables cost-effective scaling of LLM deployments by reducing inference latency and increasing throughput per GPU.
Highlights innovation in LLM optimization, which could drive demand for hardware and software supporting such techniques.
Demonstrates advanced techniques in speculative decoding and GPU optimization for AI inference.
Showcases progress in making AI models faster and more efficient, a key step toward broader accessibility.
- Speculative Decoding
- A technique to speed up LLM inference by predicting multiple tokens in parallel and verifying them in a single step.
- Parallel Tree Drafting
- A method in speculative decoding where multiple token drafts are generated in parallel using a tree structure for efficiency.
- Lossless Inference
- Generating output with no degradation in quality or accuracy compared to standard inference.
- TPS (Tokens Per Second)
- A metric measuring the throughput of an LLM, indicating how many tokens it can process per second.
- CUDA Graph
- A NVIDIA GPU optimization feature that captures and reuses sequences of operations for reduced overhead.
AI bias estimate: Neutral technical reporting with no evident bias; source is a research-focused Reddit post. (Automated estimate, not a definitive judgement.)
Don’t mistake chatbot intelligence for consciousness - The Economist
Biological AI models: new paradigms to leverage the languages of life - joint-research-centre.ec.europa.eu
China’s Military Says AI Can’t Replace Commanders. Xi Is Testing That - War on the Rocks
SPADE: Self-Play in Adaptive Synthetic Executable Environments
Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
AI ToolsMeta AI’s new Mac app wants you to talk to your apps
Meta released a new Mac application that lets users control apps and dictate text using voice commands powered by its Muse Spark AI model.
New White House strategy clarifies military tech priorities: undersea, outer space and AI - Breaking Defense
The White House released a new strategy prioritizing military investments in artificial intelligence, space systems and undersea technologies to counter emerging threats.
AI in an iron grip: How dictatorships use artificial intelligence to strengthen their rule - theins.press
A new report examines how authoritarian governments deploy AI for surveillance, censorship, and propaganda to reinforce their power.
Stripe, OpenRouter finally strike a deal - Banking Dive
Stripe and OpenRouter have partnered to integrate Stripe's payment processing with OpenRouter's AI model aggregation platform.
How one Philadelphia school is using AI to strengthen student learning, not replace teachers - CBS News
A Philadelphia school is integrating AI tools to support teachers and improve student outcomes, focusing on collaboration rather than replacement.
Exclusive-How a Texas student blew the whistle on a rogue AI hacking attempt - The Mighty 790 KFGO
A Texas student uncovered an AI-powered hacking attempt targeting local systems, prompting a swift law enforcement response.