nawfalajli

Agent VS Agent Hacker VS Developer — Testing skill for Claude Code

Testing community

Two AI agents harden your code: a Hacker agent proves each weakness with a failing test, a Developer agent fixes it — round after round until the app is production-ready.

How to install Agent VS Agent Hacker VS Developer

This entry records only its repository, not the path inside it, so there is no exact command to give. Open nawfalajli/Agent-vs-Agent-Hacker-vs-Developer and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Agent VS Agent Hacker VS Developer does

Two AI agents harden your code: a Hacker agent proves each weakness with a failing test, a Developer agent fixes it — round after round until the app is production-ready. Zero deps, runs in Claude Code.

Alternatives in Testing

  • Adversarial Spec — A Claude Code plugin that iteratively refines product specifications by debating between multiple LLMs until a 514 ★
  • Test Fixing — Detect failing tests and propose patches or fixes 475 ★
  • Production Grade — Production Grade — route any build, feature, review, test, harden, ship, or docs request through the 14-agent 175 ★

README

Agent vs Agent: Hacker & Developer

**An adversarial security automation for code you own.** A Hacker agent audits your repository and *proves* each weakness with a failing test. A Developer agent fixes the code until that test — and the whole suite — pass. Round after round, until the app is **production-ready**.

![Node](https://img.shields.io/badge/node-%3E%3D22-339933?logo=node.js&logoColor=white) ![Dependencies](https://img.shields.io/badge/dependencies-0-brightgreen) ![Runtime](https://img.shields.io/badge/runtime-Claude%20Code-d97757) ![Referee](https://img.shields.io/badge/referee-Jev%20(TypeSafe)-7c3aed) ![Tests](https://img.shields.io/badge/tests-node%3Atest-informational) ![License](https://img.shields.io/badge/license-MIT-black)


Two agents, one codebase, a refereed duel you can read like a chat log:

  • 🔴 Hacker — hunts for a weakness, states the challenge, and proves it by writing a regression test that fails on today's code.
  • 🔵 Developer — reads the challenge, fixes the application code until that test (and the full suite) pass.
  • ⚖️ Referee — the service (plus optional Jev) runs every test independently and declares a round won only when the proof is objective.

The loop repeats until a full Hacker pass finds nothing left to break. Every fix ships as a reviewed **merge request**, and you can get a **WhatsApp** message for each one.

[!IMPORTANT] This is a **defensive** tool. It runs only against repository clones you own or are authorised to test, it never contacts a live or third-party system (tests run against your own code, not a deployed host), and it ships nothing without human review. **Authorization is on you.**


Why it's trustworthy

The design makes it hard for either agent to cheat, and hard for the service to report a bug that isn't real:

Guarantee How
A finding is real, not a guess It's only acted on once its regression test