I'm an applied AI engineer building reliable agentic systems: the orchestration, evaluation, guardrails, and production infrastructure that make AI useful beyond a demo. I bring 5+ years of production reliability engineering at Texas Instruments and Dell.
- PyCastle — an autonomous development orchestrator that takes ready GitHub issues through isolated implementation, verification, integration, and a reviewed pull request. Projects own their execution graphs and fail-closed gates; PyCastle owns the recoverable runner. Supports Codex and Claude Code on the host or in Docker.
- REVAL — a fact-aligned benchmark for political and ideological bias in LLMs. It uses counterfactual pairs and a ground-truth taxonomy instead of the usual symmetry assumption. (revalbench.com)
- judge-from-scratch — an end-to-end, explained build of a specialised LLM judge: data generation, SFT, DPO, evaluation, and vLLM serving at 170+ output tokens/s on an A100.
At Dell, I built the first OpenTelemetry export pipeline for the open-source iDRAC telemetry stack, developed LLM-assisted CVE triage that reduced security review from days to under an hour, and led an explainable assistant over 100+ Redfish APIs that won Dell's 2023 Austin LLM Hackathon.
- AI systems: Python, PyTorch, Hugging Face, vLLM, LangGraph, RAG, LLM evaluation, SFT/DPO
- Backend and infrastructure: Go, FastAPI, PostgreSQL, Docker, Kubernetes, GitHub Actions
- Reliability: OpenTelemetry, Splunk, Grafana, production telemetry pipelines
I'm open to UK opportunities in applied AI, agentic AI, forward-deployed engineering, and ML platform engineering.
LinkedIn · krishnakartik1@gmail.com
Off the clock, I drum in a local band and argue about setlists.





