Discrete-event simulator of LLM inference scheduling policies, comparing FCFS, continuous batching, priority, SLO-aware, and chunked prefill under memory pressure and latency SLOs.
-
Updated
Jul 9, 2026 - C++
Discrete-event simulator of LLM inference scheduling policies, comparing FCFS, continuous batching, priority, SLO-aware, and chunked prefill under memory pressure and latency SLOs.
Lightweight edge inference runtime scheduler for priority/deadline-aware multi-task control, bounded queues, adaptive load shedding, ONNX/TensorRT-backed workers, and Jetson telemetry.
Add a description, image, and links to the inference-scheduler topic page so that developers can more easily learn about it.
To associate your repository with the inference-scheduler topic, visit your repo's landing page and select "manage topics."