Warm up once.
Serve in seconds.
A new vLLM replica spends minutes compiling kernels, autotuning and capturing CUDA graphs before its first token. warmctl does that once, snapshots the warm engine and restores it on any matching GPU host in seconds, at full serving speed.
A new replica repeats the same work. Every time.
Before a vLLM replica serves its first token, it loads and shards the weights, compiles the model, JIT-compiles and autotunes kernels for every shape, captures CUDA graphs and warms up. For a large MoE on 8 GPUs that takes more than ten minutes.
Fast storage only shortens the first step: compilers and autotuners don't care where the bytes came from. warmctl does all of that once, at snapshot time, and every replica after that starts from the warm result.
- start processes, NCCL, build model15%
- load weights7%
- compile kernels, capture CUDA graphs39%
- autotune GEMM kernels, 967 shapes17%
- warm up, capture more graphs23%
Cold start and restore measured on the same node with the same image and flags. The phase split comes from a separate cold start log of the same model on 8×H100 (673 s, weights already in page cache). Restore: CRIU 8.2 s, GPU state 10.5 s, weights 13.8 s, plus staging and admission. The throughput lead comes partly from the warmup grid the snapshot carries; stock is measured right after its own cold start.
Scale to zero. Scale to many.
The warm engine lives in the snapshot, not on a node. When traffic stops, release every GPU. When it comes back, restore as many replicas as you need, each on its own node and each serving in under a minute. No warm spares kept just in case.
From zero to serving in under 60 seconds.
Per replica, snapshot to serving with a warm page cache: GLM-5.3-Flash 37 s, Qwen3-32B up to 21 s, gpt-oss-120b up to 16 s. Replicas restore independently, so fan-out is bounded by how fast your storage serves the snapshot. Today each replica is one warmctl restore; an autoscaling operator is on the roadmap.
Freeze once. Ship anywhere. Thaw in seconds.
A snapshot holds the whole process tree of a running engine, the GPU state of every worker and the processed weights of every rank. Weights travel separately from the warm state, so a restore is never a disguised cold start.
Freeze
warmctl starts vLLM in a managed container, warms it over every declared shape, puts it to sleep and checkpoints the processes with CRIU and the GPUs with cuda-checkpoint. Weights are exported straight from vLLM's sleep-mode allocator.
$ warmctl snapshot \ --profile glm53-flash-prod-tp8 \ --model-dir /models/glm \ --gpus 0-7 --output /snaps/glm
Ship
A snapshot is a directory with a manifest, published atomically, or an artifact in any OCI registry. verify --deep hashes every file before you trust it.
$ warmctl push /snaps/glm \ oci://ghcr.io/you/glm:h100-tp8
Thaw
warmctl checks the driver and GPUs, maps the process tree back, re-attaches the GPUs, reloads weights to the same addresses, runs canaries and only then opens an OpenAI-compatible API.
$ warmctl restore /snaps/glm \ --gpus 0-7 --listen 0.0.0.0:8000 serving started 37.3s
A restore counts only if it answers like the donor did.
“Zero degradation” without checks is marketing. Every restore runs the same checks before it admits a single request.
Same addresses
CUDA graphs replay raw device pointers. The GPU address layout, communicators and every buffer come back exactly where they were.
Verified weights
Weights are checked by SHA-256 while staged and by fingerprints on the GPU, before any traffic.
Canary against the donor
The restored engine must return what the donor returned, token for token. Large models like GLM-5.3 run the canary bitwise.
Fast collectives survive
Custom, FlashInfer and symmetric-memory all-reduce and NCCL NVLS are saved and re-created at the same addresses through NVIDIA cuInterpose.
Restore to serving.
17 models proven on H100 and L40S so far, most of them from a YAML profile alone, with no code change.
| Model | Type | GPUs | Restore |
|---|---|---|---|
| GLM-5.3-Flash | MoE · FP8 · MTP | 8×H100 | 37 s |
| Qwen3-32B | dense · BF16 | 4–8×H100 | 14–21 s |
| gpt-oss-120b | MoE · MXFP4 | 2–4×H100 | 10–16 s |
| Qwen3-30B-A3B | MoE · BF16 | 1×H100 | 17–20 s |
| Gemma 4 26B-A4B | MoE · BF16 | 1×H100 | 14 s |
| Nemotron 3.5 Lightning | Mamba MoE · FP8 | 1×H100 | 11 s |
| T-pro 2.1 | dense 32B · BF16 | 2×L40S | 14 s |
| Qwen3-8B | dense · BF16 | 1×H100 | 7–10 s |
Warm page cache. A range covers every run of that model across its listed GPU counts.
One binary. Eight commands.
A static Go binary with its vLLM extension built in. No sidecar repositories, no hand-written host descriptions.
- snapshot
- bake a warmed engine into a snapshot
- restore
- bring a snapshot back to serving
- push · pull · ls
- move snapshots through any OCI registry
- verify
- check a snapshot against its manifest; --deep hashes every file
- stop
- stop a restored instance
- gc
- remove blobs no recorded snapshot uses
- doctor
- check this host
- profile
- list, check, show a profile or print its vLLM argv
Any model vLLM serves with sleep mode gets a profile in a few lines. vllm: takes any vllm serve flag; image, warmup grid and canary policy come from defaults.
schema: warmctl.profile/v1 id: llama31-8b-tp1 model: {id: meta-llama/Llama-3.1-8B-Instruct, revision: 0e9e39f2} parallel: {tp: 1} hardware: [h100-80gb-sxm] vllm: {max-model-len: 8192, kv-cache-memory-bytes: 21474836480}
Linux x86-64 · Docker + NVIDIA runtime · driver ≥ 580 · H100 SXM, L40S
Stop paying GPUs to warm up.
warmctl is in beta. Tell us which models you serve and on which GPUs, and we will profile them first.