Fast model suspend/restore for vLLM sleep mode — snapshot weights to disk and swap models on one GPU in seconds, not minutes.
-
Updated
Aug 6, 2026 - Python
Fast model suspend/restore for vLLM sleep mode — snapshot weights to disk and swap models on one GPU in seconds, not minutes.
SwapOS: frontier-class agentic AI pipelines on commodity hardware by swapping specialized small models (Context Capsule runtime)
llama-swap alternative in Rust — OpenAI-compatible GGUF runtime with VRAM-pressure eviction, context auto-fallback, and usage tracking.
Heuristic signals on whether an OpenAI-/Anthropic-compatible endpoint serves the model it claims — catch model-swapping, quantization & silent context truncation. Zero-dependency Python CLI. Signals, not proof.
Add a description, image, and links to the model-swapping topic page so that developers can more easily learn about it.
To associate your repository with the model-swapping topic, visit your repo's landing page and select "manage topics."