All work

nanoserve

An LLM inference server written from scratch, with no vLLM or TGI underneath. Built to find out where serving time actually goes.

When
Jun–Jul 2026
What
Self-led project
Stack
PyTorch, Triton, CUDA
Links
Code on GitHub

Why it exists

One request through a small model is fast. Serving is where the systems work is: hundreds of requests arrive with very different prompt and output lengths, GPU memory is finite, and you are judged on the slowest 1%, not the average.

Three things dominate. Memory is the constraint, not compute. The batch must never go stale, because a static batch of 32 spends most of its life as a batch of one while it waits for the longest sequence. And smaller weights buy more room for concurrent requests.

What I built

  • Paged KV cache. A block allocator and per-sequence block tables, so wasted memory is bounded to less than one block per sequence.
  • Continuous batching. A scheduler that re-decides the batch every step, with a token budget, chunked prefill and preemption by recompute.
  • My own forward pass for Qwen2 against the paged cache. transformers supplies only the tokenizer and the raw weights.
  • A fused Triton kernel for paged-attention decode, using online softmax over the block table.
  • Weight-only quantization. INT8 per channel and group-wise INT4. INT4 hands about 700 MiB back to the KV cache.
  • A benchmark harness with realistic lognormal traffic. vLLM runs the identical workload in a separate environment as an outside yardstick.

The scheduler and block manager import no torch, so the bookkeeping bugs that hurt most in a serving engine are caught in fast CPU unit tests. The project has 92 tests.

Results

output throughput over the best static batching (61.0 to 100.4 tokens/s)
1.65×
lower p99 time to first token (60.8 s to 5.2 s)
12×
speed-up of the Triton kernel over gather + SDPA, depending on batch and context
2.1–12.5×
of decode time is launch overhead on this GPU, measured with a roofline probe
~80%
Two charts of output tokens per second against batch size. Static batching rises with batch size but stays well below the continuous batching line in both the skewed and uniform workloads.
Continuous batching against static batching at batch sizes 1 to 16. The gap is widest on skewed lengths, which is what real traffic looks like.
Left: attention time against context length for torch and Triton backends at batch 1, 8 and 32. Right: Triton speed-up against context length, falling from about 12 times at short context to about 2 to 4 times at 4096 tokens.
The kernel’s speed-up shrinks with context at batch 1. That contradicted my prediction, and the raw timings explain why: at small batch the torch path is bound by kernel launches, not by memory traffic.

p99 end-to-end latency fell from 110.8 s to 65.4 s. The finding that decode is mostly launch-overhead bound, not memory-bandwidth bound, inverts the usual advice and drove the optimisations that followed.

What the numbers don’t say

Everything is measured on one laptop GPU (RTX 4060, 8 GB) with one small model (Qwen2.5-0.5B-Instruct, fp16). The first four attempts at one measurement gave wrong but plausible numbers, from GPU clock ramp-up to double-counted kernels. The report keeps those bugs and their fixes on record.