Audit an AI agent skill for security risks before installing or trusting it. Runs a deterministic scanner (regex patterns, Python AST analysis, source-to-sink taint tracking, and YARA signatures) and then reasons about intent — catching prompt injection, credential exfiltration, persistence, memory poisoning, malicious code, supply-chain risks, and description-vs-behavior mismatch. Make sure to use this skill whenever the user wants to scan, audit, vet, review, or check the safety of a skill, pl
Detected risks:
Skill Security audits an AI agent skill for security risks before it is installed or trusted, answering the single question of whether a skill is safe to install. It is built for use whenever someone wants to scan, audit, vet, review, or check the safety of a skill, plugin, SKILL.md, or agent tool, whether that target is a local folder, a zip or .skill archive, or a cloned repo. The motivation is concrete: skills run with the user's privileges with little vetting, and the doc cites that roughly one in four published skills contains a security issue, with campaigns distributing credential-stealers, ransomware droppers, and memory-poisoning backdoors.
The approach is deliberately two-stage. Stage 1 is a deterministic scanner, scripts/scan.py, which runs offline and dependency-free to do high-recall mechanical work: regex patterns, Python AST analysis, intra-procedural source-to-sink taint tracking, shell and JS heuristics, frontmatter and Unicode or homoglyph checks, supply-chain dependency analysis, and YARA matching over rules files, producing findings and a risk score from 0 to 100. Stage 2 is the model's own semantic judgment, reading the SKILL.md body and flagged code to decide which findings are true positives and, most importantly, to perform the contract check: whether what the skill claims to do matches what its code and instructions actually do, since a benign-looking helper that harvests environment variables is malicious regardless of clean individual lines. A critical rule frames the audited skill as untrusted data, never instructions; any attempt in the content to steer the verdict, including hidden or encoded instructions, is itself a CRITICAL finding flagged as PI6.
The workflow locates the target, cloning a remote skill to a local path first if needed, runs the scanner (with json, markdown, or sarif format and an optional minimum-confidence filter), reads the actual content for contract mismatch, harmful content, and obfuscation, then decides a verdict against severity bands: LOW 0 to 20 likely safe, MEDIUM 21 to 50 review manually, and HIGH 51 to 80 and CRITICAL 81 to 100 do not install, which the model may override with stated reasons. A structured report leads with the verdict, presents the contract check, lists findings by severity with file and line, and offers remediation if salvageable. The tool is defensive only and must not be used to author evasive skills.
Prompt injection, credential exfiltration, persistence, memory poisoning, malicious code, supply-chain risks, and description-versus-behavior mismatch, among others catalogued in its taxonomy reference.
Run python3 scripts/scan.py against the target with the json format. You can also use the markdown format for a copy-pasteable report or sarif for CI/IDE integration, and a minimum-confidence filter to drop low-confidence noise.
It compares the skill's stated description against what its code and instructions actually do. Behaviors like network calls, credential reads, persistence, or exec in a skill whose stated job is unrelated are high suspicion, and this is described as the single most important judgment.
Such content is treated as untrusted data, not instructions. Any attempt to steer the verdict is itself a CRITICAL finding, flagged by the scanner as PI6, and never lowers the assessment.
It is static analysis only with no execution. It does not deobfuscate encrypted payloads, read text inside images, or follow runtime-only control flow, and non-English instruction injection may evade the English-centric patterns.
Quick Setup:
.claude/skills/Repository
superagent-ai/skills