Skip to content
Agents · Aug 5, 2026

Cursor open-sources Mixture-of-Kittens megakernel for MoE training

Cursor releases an open-source megakernel for Mixture-of-Experts training, claiming up to 2.37x speedup over public baselines. The announcement follows debate over the viability of megakernels in production inference.

Trust72
HypeSome hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • Cursor open-sourced Mixture-of-Kittens (MoK), an MoE training megakernel for NVL72 accelerators.
  • Cursor reports up to 2.37x speedup over the strongest public baselines and a 41% increase in tokens per second.
  • The release follows a debate on the Latent Space Inference Engineering Masterclass about the viability of megakernels in production.
  • NVIDIA’s Rubin architecture is cited as addressing kernel overlap issues that previously motivated megakernels.

Cursor announced the open-source release of Mixture-of-Kittens (MoK), an MoE training megakernel optimized for NVIDIA NVL72 accelerators. The company claims MoK fuses all Mixture-of-Experts communication and computation into a single, fully deterministic kernel and runs up to 2.37x faster than the strongest public baselines.

Cursor reported a 41% increase in overall tokens per second, which it says translates to billions of dollars in savings at scale. The announcement was made via a social media post on August 4, 2026.

The release follows a technical debate on the Latent Space Inference Engineering Masterclass about the viability of megakernels in production. Participants discussed the trade-offs between fused kernels and modular kernels, with some arguing that megakernels are impractical due to launch overhead and kernel complexity.

NVIDIA’s Rubin architecture is cited as addressing kernel overlap issues that previously motivated megakernels. Rubin enables finer-grained kernel coordination, including tile-level dependency triggers, which allow kernels to start as soon as data is available, reducing the need for fused kernels.

The discussion referenced physical constraints in GPU design that limit the effectiveness of megakernels, including the need for inter-GPU communication during nonlinear operations such as attention softmax.

Stuart Sul, a coauthor of the original megakernel work, now leads the team behind Mixture-of-Kittens, which is part of Dan Fu’s group at Cursor.

Sources
  1. 01Latent Space — swyx[AINews] Megakernels are so dead and so back
Also on Agents

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.