All notes

My GPU was waiting, not working

Why LLM decode on my RTX 4060 is limited by kernel launches, not memory bandwidth, and what I changed once I knew.

The standard advice about LLM inference is that decoding is limited by memory bandwidth. Every new token means reading all of the model’s weights from GPU memory, so the speed limit is how fast you can stream bytes. I built nanoserve, an inference server written from scratch, expecting to optimise for exactly that. The first careful measurement said I was optimising the wrong thing.

The number that didn’t fit

Qwen2.5-0.5B in fp16 is 943 MiB of weights, and my laptop’s RTX 4060 sustains about 210 to 226 GB/s on a plain memory copy. If memory were the limit, one decode step at batch size 1 should take about 4.5 ms.

MeasureWindows 11WSL2 Ubuntu
Memory-bandwidth floor per step~4.5 ms~4.5 ms
Actual batch-1 step (eager transformers)23.7 ms16.6 ms
Cost of launching one small kernel7.8–15 µs6.8–9.2 µs
GPU utilisation during decode18%22%

The step was four to five times slower than the bandwidth floor. Profiling showed why: the GPU was busy for only about a fifth of each step. The rest of the time it sat idle while the CPU issued the next of roughly 1,000 to 1,350 small kernels.

The tell: batching was almost free

If most of each step is a fixed launch cost, a bigger batch should cost almost nothing extra. That is exactly what happened. Step time went from 28 ms at batch 1 to 34 ms at batch 32, so throughput went from 35.5 to 941.8 tokens per second: a 26× gain for 20% more time per step.

I blamed Windows first, and I was wrong

My first theory was the Windows display driver model (WDDM), which routes kernel launches through the operating system’s scheduler. I predicted Linux would be much faster. The measurements didn’t back that up:

  • Two consecutive runs of the same Linux setup measured the small-kernel cost at 6.8 µs and 9.2 µs. That 35% spread overlaps the Windows range.
  • The two setups also ran different PyTorch builds, which changed the number of device ops per step from 1,350 to 1,014 for identical model code. That is a library difference, not an OS one.
  • Utilisation was 18% against 22%. Both were overhead-bound.

On this hardware, the number of kernels is what moves decode time. The operating system is not.

What I changed because of it

Once the bottleneck was clear, every change that paid off was one that cut the number of ops per step:

ChangeWhy it mattered here
RoPE cos/sin computed once per stepPositions are the same in all 24 layers; the old code built 48 identical tensors per step
Fused RMSNormReplaced an 8-kernel hand-written norm used at 48 places per step
Grouped-query attention inside SDPAStopped materialising a 7× expanded copy of K and V
One pinned host-to-device copy for step metadataWas six separate small transfers per step

Together with a faster way of mapping tokens to cache slots, these took the batch-1 step from 33.7 ms to 28.2 ms, 16% faster.

Then the fused Triton paged-attention kernel replaced about ten ops per layer with one. Device ops per step fell from 1,014 to 682, GPU utilisation at batch 32 rose from 23% to 36%, and nanoserve clearly beat eager transformers for the first time: 1.23× at batch 8.

Left: attention time against context length for the torch and Triton backends. Right: the Triton kernel's speed-up, falling from about 12 times at short context to about 2 to 4 times at 4096 tokens.
The kernel’s speed-up shrinks as context grows, which surprised me. At batch 1 the torch path takes about the same time at every context length, because it is paying for ten kernel launches, not for the data.

Four measurements that lied first

Before any of these numbers could be trusted, four bugs in the measurement itself had to go. Each one produced a plausible-looking wrong answer.

  • Clock ramp. The GPU idles at 210 MHz and boosts to 2,595 MHz. Whatever ran first was measured cold, which once produced “batch 32 is faster than batch 1” and a fake 3.4× win. Now there are 8 seconds of sustained work before measuring, repeated before every row.
  • Double-counted kernels. The profiler attributes time to both an op and its kernel. Summing both reported 166% GPU utilisation.
  • Mismatched runs. Busy time from a profiled run was being divided by wall time from an unprofiled one.
  • Noise. Every configuration is now timed twice and flagged if the runs disagree by more than 15%.

Where this leaves nanoserve

Against vLLM in its strongest setup (torch.compile plus CUDA graphs), on the same GPU and the same requests, the honest comparison is:

ServerOutput tokens/svs vLLM
vLLM, compiled with CUDA graphs2,103.51.00×
nanoserve, continuous batching + Triton663.00.32×
Static batching, best batch size236.90.11×

The biggest part of that gap is the same story. vLLM captures its decode step as CUDA graphs once and replays them, while nanoserve re-issues around 680 kernels every step from Python. So the next change is CUDA graphs for decode. The shapes are static, so the step can be captured once and replayed. That attacks the time spent waiting, which is still most of each step, instead of the time spent working.

The takeaway

Measure the shape of the bottleneck on your own hardware before optimising for it. “Decode is memory-bound” may well hold for a large model on a datacenter GPU. For a 0.5B model on a laptop GPU, the real step took four to five times longer than the memory floor, and the gap was all launch overhead.