I find numbers that are wrong while nothing raises: eval scores, agent usage, cost accounting.
Twenty merged across twelve organisations, eight of them in the UK AI Security Institute's inspect_evals. Open to senior eval / agent-infrastructure roles, remote or Bengaluru. arthi1805@gmail.com
wrong-numbers : findings across the LLM eval, tracing and cost ecosystem, classified into thirteen recurring defect shapes. Twenty merged upstream after human review; eight shipped in the UK AI Security Institute's inspect_evals. Every finding links to a pull request, most with a test that fails on main. https://github.com/arthi-arumugam-git/wrong-numbers
whatbroke : diff an agent's behaviour between two runs of the same task. Tool calls, arguments, costs, outputs. https://github.com/arthi-arumugam-git/whatbroke




