ai-ml-security
AI/ML security playbook. Use when assessing model supply chain attacks (pickle RCE, poisoned weights), adversarial examples, model poisoning, model stealing, data privacy attacks (membership inference, model inversion), and autonomous agent security risks.
Security Assessment
Detected risks:
About ai-ml-security
An AI/ML security playbook covering, from an attacker's perspective, the ways machine-learning systems can be compromised, intended for security assessments of ML pipelines, model artifacts, and deployed models. It opens with model supply-chain risk, emphasizing that Python's pickle executes arbitrary code on deserialization and that PyTorch .pt/.pth files use pickle by default, making a downloaded model a code-execution vector. A format-risk table ranks .pt, .pkl, .joblib, and pickle-enabled .npy as dangerous while flagging .safetensors and .onnx as safe, and it details Hugging Face model poisoning (backdoored weights, malicious tokenizer configs, trust_remote_code) and dependency confusion in ML requirements.
Beyond the supply chain, it surveys adversarial examples with a taxonomy by attacker knowledge (white-box, transfer, query-based black-box, physical-world) and specific methods, namely FGSM, PGD, and Carlini and Wagner, plus physical-world attacks like adversarial patches, glasses, and stop-sign stickers. Model poisoning coverage includes training-data poisoning with backdoor triggers, label-flipping strategies, and gradient manipulation in federated learning, along with corresponding defenses such as robust aggregation (Krum, trimmed mean, median) and differential privacy. Model stealing is treated through query-based extraction, side-channel leakage (timing, confidence scores, top-K probabilities, power), and knowledge distillation from a black-box teacher.
Use this when threat-modeling or testing systems that load third-party models, expose inference APIs, train on external data, or run autonomous agents. It also covers data-privacy attacks (membership inference, model inversion, and gradient leakage) and LLM-specific and agent security concerns, and routes to companion skills for LLM prompt injection, insecure deserialization, and dependency confusion for deeper treatment of those specific areas.
FAQ
Why are PyTorch model files considered dangerous?
PyTorch .pt/.pth files use Python pickle by default, and pickle executes arbitrary code during deserialization, so loading an untrusted model with torch.load can run attacker code. The skill recommends weights_only=True (PyTorch 2.0+) or safetensors.
Which model formats does it consider safe?
It marks .safetensors (tensor-only, no code execution) and .onnx (graph definition only) as safe and preferred, versus critical or high risk for .pt, .pkl, .joblib, and pickle-enabled .npy.
What adversarial-example methods are covered?
FGSM (single-step), PGD (iterative), and Carlini and Wagner (optimization-based minimal perturbation), plus physical-world attacks such as adversarial patches, glasses, and stop-sign stickers.
Does it address federated learning?
Yes. It describes gradient manipulation by malicious clients (scaled, backdoor, and sign-flip gradients) and defenses like robust aggregation (Krum, trimmed mean, median), anomaly detection, and differential privacy.
How does model extraction work according to the skill?
You query a target model's API with diverse inputs, collect input/output pairs, and train a surrogate. It notes that roughly 10,000-100,000 queries often suffice for image classifiers, and covers side channels and knowledge distillation.
Install ai-ml-security
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
yaklang/hack-skills