Inference Systems Engineer · LLM Serving & Distributed Scheduling · Cloud Infrastructure
Systems Engineer with 7+ years of experience architecting and operating large-scale distributed systems at Microsoft and Amazon. Currently focused on ML-driven scheduling architectures that optimize LLM inference efficiency—spanning request admission, KV cache management, and cluster resource allocation.
Bridging machine learning research with production-grade systems engineering. Advocate for internal developer tooling, fault-tolerant infrastructure, and measurable performance optimization at scale.
| Project | Description & Impact |
|---|---|
| Clairvoyant Scheduler 📄 arXiv:2606.07248 |
Go-based sidecar proxy eliminating Head-of-Line blocking in LLM inference via ML-driven Shortest Job First scheduling. Achieves 70–76% P50 latency reduction under burst traffic without backend modifications. Target: NeurIPS 2026 Workshop → MLSys 2027 |
| ACO Sentinel | Native Kubernetes scheduler plugin combining predictive Ant Colony Optimization with trust-weighted telemetry, gRPC sidecar communication, and circuit breaker failover. <0.98ms P99 latency at 1,250 pods/sec with 46.3% cost reduction. Target: EuroSys 2027 |
| ACO: Adaptive Compute Orchestrator | Predictive job scheduler for heterogeneous compute environments using ACO with LSTM-based spike prediction and intent-aware routing. <10ms latency, 95%+ SLA adherence across 202 test scenarios. Submitted: HiPC 2026 |
| ServiceScope v2 | AI-native blast-radius analysis tool for Python microservices combining deterministic AST parsing with local LLM inference. Processes 190 files/sec with 0% inference failure and zero external API dependencies. |
| Aether Control | Enterprise LLM serving platform with integrated control plane and post-training pipeline. Built on vLLM, GRPO, and Kubernetes for end-to-end model lifecycle management. |
| vLLM Contributor | PR #41952 (Under Review): Fixed preemption ordering in PriorityRequestQueue to reduce KV cache recompute overhead and improve scheduler efficiency. |
| Domain | Technologies |
|---|---|
| Programming Languages | Python, Go, Java, Bash Currently expanding: C++, CUDA |
| Infrastructure & Orchestration | Kubernetes, Docker, Azure, AWS, Terraform CI/CD: GitHub Actions, Azure DevOps, Jenkins |
| Data & Messaging | Kafka, RabbitMQ, Azure Service Bus, Redis Databases: PostgreSQL, Cosmos DB, Neo4j |
| Observability & Reliability | Prometheus, Grafana, Azure Monitor, ELK Stack BCDR Workflows, Distributed Tracing |
- LLM Inference Optimization: Request scheduling, admission control, and KV cache management under memory and thermal constraints
- ML for Systems: Predictive modeling (LSTM, Ant Colony Optimization) for cluster resource allocation and load balancing
- Distributed Systems Engineering: Fault-tolerant, high-throughput infrastructure at billions-of-events scale
Open to discussions on LLM inference systems, distributed architecture, platform engineering, and open-source collaboration.


