Back to Skills

ai-evals

Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

1,176stars149forksUpdated 7/25/2026

Security Assessment

Safe(96/100)
Security Score96/100

About ai-evals

AI Evals is a strategy and knowledge skill that helps teams build robust infrastructure for measuring, monitoring, and iterating on AI product quality using human, code-based, and LLM-as-a-judge methodologies. It addresses the problem of teams relying on subjective 'vibe checks' to assess AI features, replacing them with systematic, empirical measurement so teams can iterate on prompts and models with a trusted feedback signal.

The skill guides users through four activities: identifying failure modes via error analysis on real traces, selecting the right mix of eval methods for a use case, building gold-standard reference datasets as ground truth, and operationalizing evaluations into a CI/CD pipeline for continuous improvement. It is built from insights and quotes attributed to eleven guests across Lenny's Podcast and Newsletter (including Brendan Foody, Edwin Chen, and Hamel Husain & Shreya Shankar) and codifies core principles such as automating measurement of a company's core value chain, prioritizing subjective excellence in data quality, eliminating vibe checks, and structuring judge-LLM logic. It supplies named templates and frameworks including an LLM-as-a-Judge playbook, a three-approaches decision framework, a four-part eval-prompt formula, open and axial coding for error analysis, a RAG evaluation framework separating retriever and generator, and an AI eval improvement flywheel, with supporting reference files for guest insights and artifacts.

It is aimed at product managers, AI engineers, and teams building or operating AI-powered products who need to move from ad-hoc quality judgments to systematic evaluation. It is a curated knowledge-and-templates skill and performs no code execution or system access itself.

FAQ

What evaluation methodologies does the skill cover?

It covers three approaches: human evaluation, code-based checks, and LLM-as-a-judge, and provides a decision framework for choosing the right mix based on your specific use case.

How does it recommend getting started with evals?

It recommends error analysis on real traces first: opening an observability tool, reviewing logs of real customer interactions, and documenting where the AI behaves unexpectedly before building any automated tests.

What frameworks and templates are included?

Named frameworks include an LLM-as-a-Judge playbook, a three-eval-approaches decision framework, a four-part eval-prompt formula, open and axial coding for error analysis, a retriever-plus-generator RAG evaluation framework, and an AI eval improvement flywheel.

Where does the content come from?

It synthesizes insights and quotes from eleven guests and posts across Lenny's Podcast and Newsletter, and ships reference files of guest insights and extracted artifacts alongside the main SKILL.md.

All Files

3 files
references/guest-insights.md29.1 KB
View
SKILL.md6.5 KB
View
references/artifacts.md45.7 KB
View

Install ai-evals

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill