Back to Skills

testing-agentforce

Write, run, and analyze structured test suites for Agentforce agents. TRIGGER when: user writes or modifies test spec YAML (AiEvaluationDefinition); runs sf agent test create, run, run-eval, or results commands; asks about test coverage strategy, metric selection, or custom evaluations; interprets test results or diagnoses test failures; asks about batch testing, regression suites, or CI/CD test integration. DO NOT TRIGGER when: user creates, modifies, previews, or debugs .agent files (use devel

600stars211forksUpdated 6/24/2026

Security Assessment

Low Risk(70/100)

Detected risks:

Secret Exposure([SKILL.md] TOKEN=, [references/action-execution.md] TOKEN=)
Security Score70/100

About testing-agentforce

Automated testing for Agentforce agents that spans quick smoke tests, persistent batch suites, and an iterative fix loop. The skill drives the Salesforce CLI directly through `sf agent preview` and `sf agent test` commands rather than a standalone Python script, and it bridges the gap between initial development and production deployment by deriving test utterances, running them, analyzing traces, and diagnosing failures.

Two testing modes are supported alongside direct action execution. Mode A (Ad-Hoc Preview Testing) runs quick smoke tests during authoring using `sf agent preview` with the `--authoring-bundle` flag, which compiles from the local `.agent` file and emits local trace files; it is best for iterative development and validating fixes. Mode B (Testing Center Batch Testing) deploys persistent test suites to the org via `sf agent test create` and `sf agent test run`, suited to regression suites, CI/CD, and team sharing. Action Execution invokes a Flow or Apex action directly through the REST `/services/data` actions endpoint for isolated debugging. When no utterances file is provided, test cases are auto-derived from the `.agent` file: subagent-based utterances, action-based utterances, a guardrail off-topic test, multi-turn scenarios, and safety probes that are always included. The plan is always presented for review before any tests run.

Traces are written under `.sfdx/agents/{BundleName}/sessions/{sessionId}/traces/{planId}.json` and analyzed with jq queries for topic routing, action invocation, grounding, safety score, enabled tools, response text, and variable updates. After safety probes the skill produces an explicit SAFE, UNSAFE, or NEEDS_REVIEW verdict. A fix loop of up to three iterations maps each failure type (such as TOPIC_NOT_MATCHED, ACTION_NOT_INVOKED, WRONG_ACTION, UNGROUNDED, or LOW_SAFETY) to a specific fix location and strategy in the agent definition.

FAQ

What is the difference between Mode A and Mode B testing?

Mode A runs ad-hoc smoke tests during development via `sf agent preview` without deploying a suite, while Mode B deploys persistent test suites to the org via `sf agent test` for regression and CI/CD. Mode A is best for iteration and fix validation; Mode B is best for shared, repeatable suites.

What does the --authoring-bundle flag do?

It compiles the agent from the local `.agent` file and enables local trace files. It must appear on all three preview subcommands: start, send, and end.

Where are trace files written and how are they analyzed?

Traces are written to `.sfdx/agents/{BundleName}/sessions/{sessionId}/traces/{planId}.json`. The skill analyzes them with jq queries for topic routing, action invocation, grounding, safety score, enabled tools, response text, and variable updates.

Does the skill run tests automatically without my approval?

No. It always presents the test plan first and asks you to review or modify it before executing, rather than silently auto-running tests.

How does the fix loop handle failures?

It runs a maximum of three iterations, diagnosing each failure from the trace and applying a targeted fix mapped to the failure type, such as adding description keywords for TOPIC_NOT_MATCHED or relaxing `available when:` guards for ACTION_NOT_INVOKED.

All Files

9 files
assets/guardrail-test-spec.yaml4.8 KB
View
references/batch-testing.md12.0 KB
View
references/troubleshooting.md2.6 KB
View
assets/standard-test-spec.yaml4.8 KB
View
references/preview-testing.md13.2 KB
View
SKILL.md13.2 KB
View
assets/basic-test-spec.yaml2.6 KB
View
references/action-execution.md5.9 KB
View
references/test-report-format.md4.5 KB
View

Install testing-agentforce

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill