Rajib-Mahmud

POC Or It Didnt Happen Web3 — Security skill for Claude Code

Security community

A Web3-only security thinking framework that supercharges any coding agent (Claude Code / Codex / opencode / DeepSeek) to audit smart contracts — generate hypotheses from every angle, then prove or re.

How to install POC Or It Didnt Happen Web3

This entry records only its repository, not the path inside it, so there is no exact command to give. Open Rajib-Mahmud/poc-or-it-didnt-happen-web3 and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What POC Or It Didnt Happen Web3 does

A Web3-only security thinking framework that supercharges any coding agent (Claude Code / Codex / opencode / DeepSeek) to audit smart contracts — generate hypotheses from every angle, then prove or reject them with runnable PoCs.

Alternatives in Security

  • /web3 Audit — Smart contract security audit using the 10-bug-class methodology 927 ★
  • OpenTag — Open-source, channel-native agent gateway for Slack 499 ★
  • Query Token Audit — Audit token security to detect scams, honeypots, and malicious contracts across BSC, Base, Solana, and Ethereu 483 ★

README

Web3 Security Thinking Framework

**A thinking framework, not an automation.** Loaded into any coding agent — **Claude Code, Codex, opencode, or DeepSeek** — it amplifies how the agent *reasons* about Web3 / smart-contract security: generate hypotheses from every angle and corner, then prove or kill each one. Precision-first — optimized to **reject non-bugs**, not just find them.

Start here: **[FRAMEWORK.md](FRAMEWORK.md)** · plan: [ROADMAP.md](ROADMAP.md) · stage spec: [PIPELINE.md](PIPELINE.md)

Use it in your agent (host-agnostic)

  • Claude Code — CLAUDE.md + the /audit skill auto-load the framework, knowledge, and memory.
  • Codex / opencode (any backend model, incl. DeepSeek) — AGENTS.md points the agent to it.
  • Any chat (DeepSeek, …) — paste FRAMEWORK.md as the system prompt, then the target code. No keys, no install.

*(`engine/` + `ingest/` are optional dev tooling — a benchmark harness and corpus ingester — not part of the framework.)*

Current status

  • M0 (foundations + eval ruler): done — knowledge files, benchmark, /audit skill.
  • M0.5 (blind validation): done — 3 targets, Precision/Recall 100%, 0 FP on a clean control (benchmarks/RESULTS.md).
  • System skeleton: in place — Web3-only engine spec (PIPELINE.md), Security Memory (memory/) with a curated seed + a real ingestion path to thousands.
  • Engine: runnable (code-complete) — python -m engine.cli selftest passes offline (dry). Real audits need a provider API key; PoC execution needs forge; thousands of cases need ingest --run.
  • Still ahead: scaled corpus (ingestion), embeddings retrieval, multi-model committee, automated PoC execution, broad benchmark. These are real, not done — see ROADMAP.md.

Structure

.
├── ROADMAP.md                         # milestone plan (+ v2 vision/architecture addendum)
├── PIPELINE.md                        # model-agnostic engine spec (the portable "intelligence")
├── README.md
├── .claude/skills