Skip to content
Main Site News Console

[AINews] Megakernel Was Completely Dead—Now It’s Back

· Latent Space Translated
播客深度访谈

Yesterday, there was a heated discussion about Megakernel on our Inference Engineering Masterclass podcast:

Megakernel is dead

Why is Megakernel useful? You spend two months writing a Kernel just to save launch overhead and solve the problem that Kernels don’t overlap well.

There used to be PDL, but then people said it wasn’t perfect because you could still get some marginal gains due to straggler CTAs. So—wait, sorry, I forgot: Rubin solves that problem. (Kernel 2 needs 10 CTAs, while Kernel 1 has already completed 7 and has 3 stragglers; Kernel 2 can now launch 7 of its CTAs first.)

Given a long enough timeline, everything eventually evens out. No serious inference provider would use a 67,000-line, manually fused forward-pass Kernel in production; teams that do this are purely conducting research.

Dead.

Two weeks ago, I appeared on @swyx’s podcast and said some things I… probably shouldn’t have said.

A lot has happened since then, and I owe everyone an apology.

I’m sorry that everything I said turned out to be correct.

a) On “Megakernel is dead”

Why is Megakernel useful? You spend two months…

For those willing to listen to the full discussion, here it is:

Ali: Fused Kernels can’t save you. For example, in tensor parallelism, half of a matrix is on one GPU and the other half is on another GPU. If the next step requires the entire matrix to perform a nonlinear operation—for example, if you need to perform softmax during attention, or calculate an exponential—then I need to have a complete row of data. So I need to know the partial results from GPU 2 and the partial results from GPU 1 before I can perform softmax in the next stage.

So even if I use fused Kernels, they still have to communicate with one another, because there are nonlinear operations within each part. As for Megakernel, to be honest, I’m very skeptical. It’s a good area of research, and intuitively and theoretically, it seems beautiful. You can eliminate a lot of launch overhead by launching only one Kernel—continually fusing and moving data, and combining everything together.

But the complexity of the Kernel itself makes it very difficult—really very difficult—to write a highly optimized Megakernel. I’m not referring to any specific company, but even companies that have developed fused Megakernels, or people I’ve spoken to who work at such companies, often end up not running these Kernels in production because TensorRT-LLM and modular Kernels launch faster: you can optimize each component independently and also run them in parallel.

A technical lead at NVIDIA once tweeted, “We’re going to unveil Rubin; here are its specifications.” The third tweet showed the relevant information. I don’t want to go into the technical details yet, and I need to read through it more carefully, but the way this GPU is designed will make Megakernel meaningless. So it seems that the entire research area is not going to continue developing.

He was referring to Kyle Kranen’s explanation of dependency triggers (a friend of this show!). This was one of the factors that previously caused pipelines to stall, making Kernel fusion necessary:

Improved Kernel overlap: Rubin supports finer-grained Kernel coordination, including tile-level dependency triggers. This means that as soon as part of an operation’s data becomes available, the Kernel responsible for executing that part of the operation can start immediately!

As discussed on the show, some physical limitations still remain unresolved. But it makes perfect sense for NVIDIA to update Rubin’s design so that it better accommodates the extreme optimizations currently underway in the Kernel space.

Stuart Sul, one of Ben Spector’s Megakernel collaborators, now leads the team behind Mixture of Kittens. Mixture of Kittens is an open-source Megakernel released by Cursor today for Mixture of Experts (MoE) training.

The name pays homage to Ben’s highly personal project ThunderKittens, which is part of Dan Fu’s research team:

We are open-sourcing Mixture-of-Kittens (MoK), our MoE training Megakernel for NVL72.

It fuses all Mixture-of-Experts communication and computation into a single fully deterministic Kernel, achieving speeds of up to 2.37× the strongest publicly available baseline.