Exploitation Validator — Data skill for Claude Code
A prompt-based pipeline for finding, validating, and proving vulnerabilities using LLM sub-agents.
How to install Exploitation Validator
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open gadievron/exploitation-validator and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Exploitation Validator does
A prompt-based pipeline for finding, validating, and proving vulnerabilities using LLM sub-agents.
Alternatives in Data
- LLM Redteam — LLM red-team corpus runner — fires categorized prompt-injection / jailbreak / system-prompt-leak / data-exfil 4.5k ★
- 06 Ink React Terminal UI — Prompt 06: Verify and Fix the Ink/React Terminal UI Pipeline 2.3k ★
- Vibesec — Helps write secure code by preventing common vulnerabilities including IDOR, XSS, SQL injection, SSRF, and wea 687 ★
README
exploitation-validator, an Exploitability Validation Skill / System
A prompt-based pipeline for finding, validating, and proving vulnerabilities using LLM sub-agents — structured to resist false positives.
Authors: Gadi Evron ([@gadievron](https://github.com/gadievron)) and Michal Kamensky ([@kamenskymic](https://github.com/kamenskymic))
Note: This system has since been enhanced and turned into a skill by John Cartwright ([@grokjc](https://github.com/grokjc)), where he combined it with his binary exploitation module for [raptor](https://github.com/gadievron/raptor).
**And if you like what I do**, check out my startup [Knostic](https://knostic.ai) where we protect coding agents/MCP/extensions/skills, etc.
What It Does
Takes a codebase, searches for vulnerabilities (e.g., command injection), validates findings aren't hallucinated, known, or by-design, and — where the environment allows executing a harmless PoC — proves real ones. A finding verified only by static dataflow is reported as `confirmed`; `exploitable` requires an observed effect (see the execution model in `shared.md`).
Stages
This pipeline is packaged as a skill — **`SKILL.md`** is the operational entry point and the source of truth for the stage table. The copies below are a human-readable mirror; if they ever diverge, `SKILL.md` wins.
| Stage | Purpose | Output |
|---|---|---|
| 0: Inventory | Build ground truth checklist of all files/functions | checklist.json |
| A: One-Shot | Quick exploitability check + PoC attempt (a success does not skip B/C) | findings.json |
| B: Process | Systematic analysis with attack trees, hypotheses, multiple paths | findings.json + working docs |
| C: Sanity | Validate LLM didn't hallucinate — mechanical fact-check only (files exist, code matches, flow is real) | validated findings.json |
| C-bis: Semantic | Is it actually a vuln? Rule out algorithm tautologies, spec-required behavior, documented design | findings.js |
Related Skills
AI Coding Exercises
Hands-on practice for AI engineering interviews: build LLM evals from scratch with the Claude API, plus drills
SQL AI Agent
A text-to-SQL AI agent with a safety-first validation layer. Natural language questions become SQL, but a stri
Digital Twin Of Yourself
Reverse-engineer how you think, talk, and make decisions — then turn it into a stress-tested AI System Prompt.
Cost Aware LLM Pipeline
Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic,
Scan Diff
Scan only the changed files in a single commit or PR/MR for newly introduced vulnerabilities (fast incremental
Gsd Submit
File a verified finding as a proper open-gsd/gsd-core issue + fix PR, following the core-contribution skill's
Related Agents
Web Pentester
Authorized offensive security testing of web applications you own or have written permission to test - finding
AI Redteam
Use PROACTIVELY for AI CTF challenges and LLM security challenges. Triggers on prompt injection, jailbreak, LL
Exploit Verifier Prosecutor
Phase 4 of the mobile-security-skill pipeline (half of the adversarial exploitation panel). Builds the stronge