Agent Skill for Evaluation-Driven Development (EDD) - guide AI evaluation strategies with the 50-40-10 rule
-
Updated
Jan 11, 2026 - Python
Agent Skill for Evaluation-Driven Development (EDD) - guide AI evaluation strategies with the 50-40-10 rule
A generic, domain-agnostic Python library for Evaluation-Driven Development (EDD): measure non-deterministic LLM output against an answer key, with before/after baselines to catch regression and drift — the agent-era analog of TDD.
Python package that generalizes the testing of extracted entities and features using LLMs and Pydantic models
Trajectory analysis framework for agent evaluation — moving beyond reward scores to behavioral explanation.
Add a description, image, and links to the evaluation-driven-development topic page so that developers can more easily learn about it.
To associate your repository with the evaluation-driven-development topic, visit your repo's landing page and select "manage topics."