AI ToolsAug 4, 2026, 6:38 PM

Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

30-second summary

Cursor Research has open-sourced Mixture-of-Kittens (MoK), a deterministic mixture-of-experts training kernel that accelerates model training by up to 2.37x on NVIDIA GB300 NVL72 racks.

TickrWire
Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks
Key takeaways
  • MoK is a deterministic mixture-of-experts training kernel that fuses communication and computation into a single operation.
  • It delivers up to 2.37x speed improvements on NVIDIA GB300 NVL72 racks compared to public baselines.
  • Requires Blackwell SM100 or SM103 GPUs, limiting accessibility to organizations with NVL72 capacity.
  • Cursor Research has open-sourced MoK to encourage broader adoption and innovation in MoE training.
Full story

Cursor Research has open-sourced Mixture-of-Kittens (MoK), a deterministic mixture-of-experts (MoE) training megakernel designed to optimize performance on NVIDIA GB300 NVL72 racks. The kernel consolidates all MoE communication and computation into a single, deterministic operation, significantly reducing overhead and improving training speed.

According to Cursor, MoK achieves up to 2.37x faster training compared to the strongest public baseline when running on GB300 NVL72 systems. This performance boost is tied to the use of Blackwell SM100 or SM103 GPUs, which are required for MoK to function. The open-source release targets developers working with large-scale MoE models, particularly those leveraging NVIDIA's latest hardware.

The move underscores Cursor's commitment to advancing efficient training methodologies for large language models. By releasing MoK under an open-source license, the company aims to foster collaboration and innovation within the AI research community.

Sponsored
Why this matters
Developers

Provides a high-performance, deterministic MoE training kernel for large-scale models.

Businesses

Enables faster model training on NVIDIA GB300 systems, reducing computational costs.

Investors

Highlights advancements in AI training efficiency, potentially increasing value in AI infrastructure.

Everyone

Demonstrates progress in optimizing AI training for next-generation hardware.

Glossary
Mixture-of-Experts (MoE)
A machine learning model architecture that uses multiple specialized sub-models (experts) and selectively activates only a subset for each input, improving efficiency.
Megakernel
A single, highly optimized kernel that combines multiple operations to reduce overhead and improve performance.
GB300 NVL72
NVIDIA's high-performance computing system designed for large-scale AI workloads, featuring Blackwell architecture GPUs.
Sources · 1
Read next
More stories
WeatherNext: AI model achieves breakthrough in forecasting cyclonesAI Research

WeatherNext: AI model achieves breakthrough in forecasting cyclones

DeepMind introduced WeatherNext, an AI system that markedly improves cyclone track and intensity predictions, extending forecast lead times by several days.

TickrWire

US Senate Commerce approves KOSA, children's AI safety bills - IAPP

The US Senate Commerce Committee has approved two bills focused on AI safety for children. The bills aim to regulate AI systems and protect children's data.

TickrWire
AI Research

Artificial intelligence enters Italy’s national security agenda - Decode39

Italy has added artificial intelligence to its national security agenda, marking a significant development in the country's approach to AI. This move is expected to have implications for the nation's defense and security strategies.

TickrWire

Powering the ballot: Why AI’s energy footprint is the ultimate midterm election issue - Route Fifty

AI’s growing energy demands are becoming a key issue in the US midterm elections, raising questions about sustainability and infrastructure.

Sponsored
TickrWire
AI Research

Accelerating Biomedical Innovation with AI through Collaborative Iteration - Wyss Institute at Harvard

The Wyss Institute at Harvard is leveraging AI to accelerate biomedical innovation through collaborative iteration. Researchers are using AI to analyze and improve medical devices and treatments.

TickrWire
Funding

DeepSeek invests $20.8 million in Unitree's Shanghai IPO - Reuters

DeepSeek has committed $20.8 million to Unitree's upcoming Shanghai IPO, signaling strong investor confidence in the robotics firm.

TickrWireAI News Intelligence

We aggregate, verify, summarise and explain the latest artificial intelligence news from open, legal sources.

Daily AI digest

Top AI stories, summarised, in your inbox each morning.

© 2026 TickrWire. Summaries and analysis are AI-generated and may contain errors.