Warm restore for vLLM

Warm up once.
Serve in seconds.

A new vLLM replica spends minutes compiling kernels, autotuning and capturing CUDA graphs before its first token. warmctl does that once, snapshots the warm engine and restores it on any matching GPU host in seconds, at full serving speed.

new replica · 8×H100
A new GLM-5.3-Flash replica, from snapshot to serving in 37 s. A cold start of the same engine takes 650 s.
Cold start

A new replica repeats the same work. Every time.

Before a vLLM replica serves its first token, it loads and shards the weights, compiles the model, JIT-compiles and autotunes kernels for every shape, captures CUDA graphs and warms up. For a large MoE on 8 GPUs that takes more than ten minutes.

Fast storage only shortens the first step: compilers and autotuners don't care where the bytes came from. warmctl does all of that once, at snapshot time, and every replica after that starts from the warm result.

GLM-5.3-Flash · FP8 · 8×H100new replica → serving
vLLM cold start
650 s
  • start processes, NCCL, build model15%
  • load weights7%
  • compile kernels, capture CUDA graphs39%
  • autotune GEMM kernels, 967 shapes17%
  • warm up, capture more graphs23%
warmctl restore
37 s
17×faster than cold start
100–136%of stock throughput after
bitwisecanary match with the donor

Cold start and restore measured on the same node with the same image and flags. The phase split comes from a separate cold start log of the same model on 8×H100 (673 s, weights already in page cache). Restore: CRIU 8.2 s, GPU state 10.5 s, weights 13.8 s, plus staging and admission. The throughput lead comes partly from the warmup grid the snapshot carries; stock is measured right after its own cold start.

Scale

Scale to zero. Scale to many.

The warm engine lives in the snapshot, not on a node. When traffic stops, release every GPU. When it comes back, restore as many replicas as you need, each on its own node and each serving in under a minute. No warm spares kept just in case.

GLM-5.3-Flash8×H100
0replicas serving
GPUs in use0
each replica37 s to serving
statescaled to zero
GLM-5.3-Flash

From zero to serving in under 60 seconds.

Per replica, snapshot to serving with a warm page cache: GLM-5.3-Flash 37 s, Qwen3-32B up to 21 s, gpt-oss-120b up to 16 s. Replicas restore independently, so fan-out is bounded by how fast your storage serves the snapshot. Today each replica is one warmctl restore; an autoscaling operator is on the roadmap.

How it works

Freeze once. Ship anywhere. Thaw in seconds.

A snapshot holds the whole process tree of a running engine, the GPU state of every worker and the processed weights of every rank. Weights travel separately from the warm state, so a restore is never a disguised cold start.

01

Freeze

warmctl starts vLLM in a managed container, warms it over every declared shape, puts it to sleep and checkpoints the processes with CRIU and the GPUs with cuda-checkpoint. Weights are exported straight from vLLM's sleep-mode allocator.

$ warmctl snapshot \
    --profile glm53-flash-prod-tp8 \
    --model-dir /models/glm \
    --gpus 0-7 --output /snaps/glm
02

Ship

A snapshot is a directory with a manifest, published atomically, or an artifact in any OCI registry. verify --deep hashes every file before you trust it.

$ warmctl push /snaps/glm \
    oci://ghcr.io/you/glm:h100-tp8
03

Thaw

warmctl checks the driver and GPUs, maps the process tree back, re-attaches the GPUs, reloads weights to the same addresses, runs canaries and only then opens an OpenAI-compatible API.

$ warmctl restore /snaps/glm \
    --gpus 0-7 --listen 0.0.0.0:8000
  serving started 37.3s
Guarantees

A restore counts only if it answers like the donor did.

“Zero degradation” without checks is marketing. Every restore runs the same checks before it admits a single request.

Same addresses

CUDA graphs replay raw device pointers. The GPU address layout, communicators and every buffer come back exactly where they were.

Verified weights

Weights are checked by SHA-256 while staged and by fingerprints on the GPU, before any traffic.

Canary against the donor

The restored engine must return what the donor returned, token for token. Large models like GLM-5.3 run the canary bitwise.

Fast collectives survive

Custom, FlashInfer and symmetric-memory all-reduce and NCCL NVLS are saved and re-created at the same addresses through NVIDIA cuInterpose.

Measured

Restore to serving.

17 models proven on H100 and L40S so far, most of them from a YAML profile alone, with no code change.

ModelTypeGPUsRestore
GLM-5.3-FlashMoE · FP8 · MTP8×H10037 s
Qwen3-32Bdense · BF164–8×H10014–21 s
gpt-oss-120bMoE · MXFP42–4×H10010–16 s
Qwen3-30B-A3BMoE · BF161×H10017–20 s
Gemma 4 26B-A4BMoE · BF161×H10014 s
Nemotron 3.5 LightningMamba MoE · FP81×H10011 s
T-pro 2.1dense 32B · BF162×L40S14 s
Qwen3-8Bdense · BF161×H1007–10 s

Warm page cache. A range covers every run of that model across its listed GPU counts.

CLI

One binary. Eight commands.

A static Go binary with its vLLM extension built in. No sidecar repositories, no hand-written host descriptions.

snapshot
bake a warmed engine into a snapshot
restore
bring a snapshot back to serving
push · pull · ls
move snapshots through any OCI registry
verify
check a snapshot against its manifest; --deep hashes every file
stop
stop a restored instance
gc
remove blobs no recorded snapshot uses
doctor
check this host
profile
list, check, show a profile or print its vLLM argv

Any model vLLM serves with sleep mode gets a profile in a few lines. vllm: takes any vllm serve flag; image, warmup grid and canary policy come from defaults.

llama31-8b-tp1.yaml
schema: warmctl.profile/v1
id: llama31-8b-tp1
model: {id: meta-llama/Llama-3.1-8B-Instruct, revision: 0e9e39f2}
parallel: {tp: 1}
hardware: [h100-80gb-sxm]
vllm: {max-model-len: 8192, kv-cache-memory-bytes: 21474836480}

Linux x86-64 · Docker + NVIDIA runtime · driver ≥ 580 · H100 SXM, L40S

Early access

Stop paying GPUs to warm up.

warmctl is in beta. Tell us which models you serve and on which GPUs, and we will profile them first.