Agent VS Agent Hacker VS Developer — Testing skill for Claude Code
Two AI agents harden your code: a Hacker agent proves each weakness with a failing test, a Developer agent fixes it — round after round until the app is production-ready.
How to install Agent VS Agent Hacker VS Developer
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open nawfalajli/Agent-vs-Agent-Hacker-vs-Developer and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Agent VS Agent Hacker VS Developer does
Two AI agents harden your code: a Hacker agent proves each weakness with a failing test, a Developer agent fixes it — round after round until the app is production-ready. Zero deps, runs in Claude Code.
Alternatives in Testing
- Adversarial Spec — A Claude Code plugin that iteratively refines product specifications by debating between multiple LLMs until a 514 ★
- Test Fixing — Detect failing tests and propose patches or fixes 475 ★
- Production Grade — Production Grade — route any build, feature, review, test, harden, ship, or docs request through the 14-agent 175 ★
README
Agent vs Agent: Hacker & Developer
**An adversarial security automation for code you own.** A Hacker agent audits your repository and *proves* each weakness with a failing test. A Developer agent fixes the code until that test — and the whole suite — pass. Round after round, until the app is **production-ready**.
   -7c3aed)  
Two agents, one codebase, a refereed duel you can read like a chat log:
- 🔴 Hacker — hunts for a weakness, states the challenge, and proves it by writing a regression test that fails on today's code.
- 🔵 Developer — reads the challenge, fixes the application code until that test (and the full suite) pass.
- ⚖️ Referee — the service (plus optional Jev) runs every test independently and declares a round won only when the proof is objective.
The loop repeats until a full Hacker pass finds nothing left to break. Every fix ships as a reviewed **merge request**, and you can get a **WhatsApp** message for each one.
[!IMPORTANT] This is a **defensive** tool. It runs only against repository clones you own or are authorised to test, it never contacts a live or third-party system (tests run against your own code, not a deployed host), and it ships nothing without human review. **Authorization is on you.**
Why it's trustworthy
The design makes it hard for either agent to cheat, and hard for the service to report a bug that isn't real:
| Guarantee | How |
|---|---|
| A finding is real, not a guess | It's only acted on once its regression test |
Related Skills
Hoofy
Hoofy — AI development companion MCP server. Persistent memory, spec-driven development, adaptive change pipel
Happysquad Loop
Run the full happysquad loop (architect → implement → test → conflict-gate → review) on a task, iterating unti
Mutants
Mutation-test changed files; strengthen tests until zero survivors
Motion Design
Design and spec motion for a screen, component, or page — picks tool, durations, easings, and writes a develop
Twin Sparrow Expert Engineer
Multi-domain engineering skill covering backend, frontend, TypeScript, debugging, architecture, performance, t
AI Review Pipeline
AI code review pipeline — review, auto-fix, test generation & HTML report in one command. Zero dependencies, 6
Related Agents
Weavie Tester
Proves a change actually works by running the real Weavie app, exercising the scenarios a PR should cover, and
Evidence Reviewer
Skeptical auditor of evidence bundles against .claude/skills/evidence-standards.md. Detects circular citations
Codex Debugger
Root-cause a failing test, crash, stack trace, or misbehaving feature by delegating the investigation to OpenA