nanoserve
An LLM inference server written from scratch, with no vLLM or TGI underneath. Built to find out where serving time actually goes.
Why it exists
One request through a small model is fast. Serving is where the systems work is: hundreds of requests arrive with very different prompt and output lengths, GPU memory is finite, and you are judged on the slowest 1%, not the average.
Three things dominate. Memory is the constraint, not compute. The batch must never go stale, because a static batch of 32 spends most of its life as a batch of one while it waits for the longest sequence. And smaller weights buy more room for concurrent requests.
What I built
- Paged KV cache. A block allocator and per-sequence block tables, so wasted memory is bounded to less than one block per sequence.
- Continuous batching. A scheduler that re-decides the batch every step, with a token budget, chunked prefill and preemption by recompute.
- My own forward pass for Qwen2 against the paged cache.
transformerssupplies only the tokenizer and the raw weights. - A fused Triton kernel for paged-attention decode, using online softmax over the block table.
- Weight-only quantization. INT8 per channel and group-wise INT4. INT4 hands about 700 MiB back to the KV cache.
- A benchmark harness with realistic lognormal traffic. vLLM runs the identical workload in a separate environment as an outside yardstick.
The scheduler and block manager import no torch, so the bookkeeping bugs that hurt most in a serving engine are caught in fast CPU unit tests. The project has 92 tests.
Results
- output throughput over the best static batching (61.0 to 100.4 tokens/s)
- 1.65×
- lower p99 time to first token (60.8 s to 5.2 s)
- 12×
- speed-up of the Triton kernel over gather + SDPA, depending on batch and context
- 2.1–12.5×
- of decode time is launch overhead on this GPU, measured with a roofline probe
- ~80%


p99 end-to-end latency fell from 110.8 s to 65.4 s. The finding that decode is mostly launch-overhead bound, not memory-bandwidth bound, inverts the usual advice and drove the optimisations that followed.
What the numbers don’t say
Everything is measured on one laptop GPU (RTX 4060, 8 GB) with one small model (Qwen2.5-0.5B-Instruct, fp16). The first four attempts at one measurement gave wrong but plausible numbers, from GPU clock ramp-up to double-counted kernels. The report keeps those bugs and their fixes on record.