Back to Skills

google-agents-cli-eval

This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an evalset", "debug eval scores", "compare eval results", or needs guidance on ADK (Agent Development Kit) evaluation methodology and the eval-fix loop. Covers eval metrics, evalset schema, LLM-as-judge, tool trajectory scoring, and common failure causes. Part of the Google ADK (Agent Development Kit) skills suite. Do NOT use for API code patterns (use google-agents-cli-adk-code), deployment (us

2,079stars246forksUpdated 5/6/2026

Security Assessment

Critical Risk(25/100)

Detected risks:

Remote Code Execution([references/builtin-tools-eval.md] eval()
Secret Exposure([references/builtin-tools-eval.md] api_key =)
Sensitive File Access([references/builtin-tools-eval.md] .env, [references/multimodal-eval.md] .env)
Security Score25/100

About google-agents-cli-eval

The `google-agents-cli-eval` skill is designed to help users systematically evaluate agents built using the Google Agent Development Kit (ADK). It addresses the need for structured assessment of agent performance by providing guidance on evaluation methodology, metrics, and iterative improvement processes. This skill allows users to run evaluation sets, debug evaluation scores, and compare results across different versions of an agent, ensuring that the agent behaves as expected in a variety of scenarios. By focusing on evaluation rather than deployment or scaffolding, it provides a targeted approach to improving agent quality and reliability within the ADK ecosystem.

Key features of this skill include support for running eval sets through the `agents-cli eval run` command, references for evaluation criteria, dynamic user simulation, and tool-specific trajectory scoring. It covers both standard and multimodal evaluation inputs, detailing schema requirements and metric limitations. Users can leverage the Eval-Fix loop methodology to iteratively diagnose low scores, adjust agent instructions or eval sets, and rerun evaluations to track improvements. Additional guidance is provided on common pitfalls, such as improper threshold tuning, flaky evals, or misalignment between expected outputs and agent behavior.

This skill is primarily intended for developers and testers working with ADK agents who need to ensure high-quality performance and correctness. Typical use cases include evaluating new agent versions, debugging failing evaluation cases, refining agent instructions, and tracking improvements across iterative fixes. It is suitable for both scaffolded projects and standalone agents where evaluation rigor is critical, making it valuable for teams aiming to systematically enhance agent capabilities and reduce unintended behavior in production scenarios.

FAQ

How do I start evaluating my ADK agent?

Install the `agents-cli` tool using `uv tool install google-agents-cli` and run `agents-cli eval run` to execute your evaluation set. If you used a scaffolded project, the necessary eval configuration and evalsets are already provided.

Can I evaluate agents that use multimodal inputs?

Yes, the skill includes guidance on multimodal evaluation, including evalset schema, built-in metric limitations, and custom evaluator patterns.

What should I do if an eval score is below the threshold?

Follow the Eval-Fix loop: diagnose the cause of the failure, adjust agent instructions, tool logic, or evalset as needed, and rerun the evaluation. Repeat until the evaluation passes.

Is this skill suitable for deployment or code generation tasks?

No, this skill focuses solely on evaluation. For deployment, use `google-agents-cli-deploy`, and for code generation patterns, use `google-agents-cli-adk-code`.

What are common mistakes to avoid during evaluation?

Avoid lowering thresholds to artificially pass cases, skipping flaky evals, or only adjusting evalsets instead of fixing agent behavior. These shortcuts can hide real issues and cost more time in the long run.

All Files

5 files
SKILL.md13.8 KB
View
references/user-simulation.md3.8 KB
View
references/builtin-tools-eval.md5.1 KB
View
references/criteria-guide.md5.0 KB
View
references/multimodal-eval.md5.1 KB
View

Install google-agents-cli-eval

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill