agent-eval
Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics
Security Assessment
About agent-eval
The agent-eval skill is a command-line tool designed to systematically compare coding agents such as Claude Code, Aider, and Codex on custom development tasks. It addresses the challenge of determining which AI coding assistant performs best in a reproducible and quantifiable manner. By automating agent evaluation, it eliminates the guesswork and subjective assessments commonly involved in choosing AI tools for coding workflows. The skill emphasizes reproducibility by using isolated git worktrees, ensuring each agent run is independent and cannot interfere with the base code repository.
Agent-eval offers several key features that enable comprehensive evaluation of coding agents. Tasks are defined declaratively using YAML, specifying the required files, the prompt for the agent, and the judge criteria for success. It supports multiple judge types, including deterministic code-based tests, pattern-based verification using grep, and even LLM-based judgment for subjective assessments. Metrics collected include pass rate, cost, execution time, and consistency across repeated runs, providing a multidimensional view of agent performance. Users can run multiple agents concurrently on the same tasks and generate comparison reports in table format to visualize results easily.
This skill is particularly useful for developers, team leads, and technical decision-makers who want data-backed insights before adopting new coding assistants. It is suited for use cases such as evaluating agent performance on existing codebases, conducting regression checks after agent updates, and selecting the most efficient tool for a team's workflow. By enabling reproducible and systematic benchmarking, agent-eval empowers teams to make informed decisions on which coding agent to rely on for production or experimental tasks.
FAQ
How do I define tasks for agent-eval?
Tasks are defined using YAML files that specify the task name, description, target files, prompt for the agent, judge criteria, and a commit hash for reproducibility.
Which coding agents are supported?
Agent-eval supports head-to-head comparison of coding agents such as Claude Code, Aider, Codex, and potentially others that can be run via the CLI interface.
Do I need Docker to run agent-eval?
No, Docker is not required. Each agent run is isolated using a separate git worktree to ensure reproducibility and prevent interference with the base repository.
What types of judges can I use to evaluate agent output?
You can use code-based deterministic judges (e.g., pytest or build commands), pattern-based judges (e.g., grep), and LLM-based judges that review implementation correctness.
How many runs per agent are recommended?
It is recommended to run at least three trials per agent to account for non-deterministic behavior and to calculate consistency metrics accurately.
Install agent-eval
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
affaan-m/everything-claude-code