Adversarial Dev — Development skill for Claude Code
GAN-inspired three-agent harness that pits a generator against an adversarial evaluator to build applications with quality gates at every step.
How to install Adversarial Dev
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open coleam00/adversarial-dev and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Adversarial Dev does
GAN-inspired three-agent harness that pits a generator against an adversarial evaluator to build applications with quality gates at every step. Built with Claude Agent SDK and Codex SDK.
Alternatives in Development
- Flow Parallel — Decompose and execute large changes, migrations, or multi-issue fixes in parallel with quality gates 2.8k ★
- Webgpu Claude Skill — A Claude skill for developing WebGPU applications with Three.js 610 ★
- Claudeforge — A CLAUDE.md Generator and Maintenance tool for for Claude Code to create high-quality CLAUDE.md instruction fi 318 ★
README
Adversarial Dev
A GAN-inspired three-agent harness that separates **planning**, **building**, and **evaluation** into distinct AI agents with distinct contexts. The evaluator's job is to **break** what the generator builds -- creating adversarial tension that drives quality far beyond what a single agent can achieve. Built with both the **Claude Agent SDK** and **Codex SDK** so you can run the same architecture on either platform.
Based on Anthropic's engineering article: [Harness Design for Long-Running Application Development](https://www.anthropic.com/engineering/harness-design-long-running-apps).
What This Demonstrates
Most AI coding agents fail on complex tasks not because the model is bad, but because nobody separated the work into specialized roles. A single agent that plans, builds, and evaluates its own work will reliably praise its own mediocre output. This is called **self-evaluation bias**, and it's the quiet killer of ambitious AI coding projects.
This project implements the fix: three agents, each with a focused job and its own context window.
| Agent | Role | Analogy |
|---|---|---|
| Planner | Expands a short prompt into a full product spec with sprints | Product manager |
| Generator | Builds one feature at a time, commits to git | Software engineer |
| Evaluator | Actively tries to break what the generator built, scores ruthlessly | Adversarial QA |
The evaluator doesn't just review code -- it's an adversary. It runs the application, probes for failures, tests edge cases the generator didn't think of, and scores each criterion on a 1-10 scale with a hard pass threshold. If any criterion fails, the sprint goes back to the generator with detailed, unforgiving feedback. The generator has to fight its way past the evaluator to advance. This adversarial pressure is what turns AI-generated code from "looks right" into "actually works."
Quick Start
Prerequisites
- Bun runtime installed
- Cla
Related Skills
Gan Build
Run a generator/evaluator build loop for implementation tasks with bounded iterations and scoring.
Loop Execution Evaluator
Verify implementation quality. Evaluate-Loop Step 4.
Nova Lancer
Star Fox-inspired 3D rail shooter for the browser, built with Claude Code × Claude Opus 5.5 (MAX) — three.js,
Opc
OPC — One Person Company. A full team in a single Claude Code skill. Adaptive agent orchestrator with 11 built
Gate Check
Execute the three mandatory quality gates for the current project.
Polymarket Bot Graveyard
An honest, step-by-step post-mortem of six autonomous Polymarket trading bots built and killed over three mont
Related Agents
Gan Evaluator
GAN Harness — Evaluator agent. Tests the live running application via Playwright, scores against rubric, and p
Harness Evaluator
Use this agent when the harness needs to evaluate a generator's output. The evaluator interacts with the runni
UI Evaluator
Adversarial UI evaluator that tests live web applications via Playwright. Use only in projects with a running