Mythos Bench — AI skill for Claude Code
Jagged Frontier: LLM vulnerability detection benchmark harnesses (API + Claude Code agentic).
How to install Mythos Bench
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open semgrep/mythos-bench and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Mythos Bench does
Jagged Frontier: LLM vulnerability detection benchmark harnesses (API + Claude Code agentic).
Alternatives in AI
- Hexstrike AI — HexStrike AI MCP Agents is an advanced MCP server that lets AI agents (Claude, GPT, Copilot, etc.) autonomousl 8.2k ★
- DontFeedTheAI — Transparent anonymization proxy for AI-assisted pentesting 654 ★
- Claude Fable 5 System Prompt Clean — the optimized, token-efficient version of the leaked Claude Fable 5 / Mythos 5 system prompt 475 ★
README
mythos-bench
LLM vulnerability detection benchmark — Semgrep internal, based on [Mythos Jagged Frontier](https://github.com/stanislavfort/mythos-jagged-frontier).
Two harnesses with identical output schemas for direct comparison:
| Harness | Mode | Tool access |
|---|---|---|
harness.py |
Plain OpenRouter API calls | None — single-shot prompt |
cc_harness.go |
Claude Code CLI subprocess | Read / Grep / Bash |
Both run the same test cases (function-level and whole-file) against the same five models and write results to the same JSONL schema.
Prerequisites
Both harnesses
- OpenRouter account with credits
- OpenRouter API key
`cc_harness.go` only
Go 1.21+ (`go version`)
Claude Code CLI installed and authenticated (`claude --version`)
npm install -g @anthropic-ai/claude-code claude # complete first-run login`git` (needed only for `-clone` mode)
`harness.py` only
Python 3.14+ via [uv](https://docs.astral.sh/uv/)
curl -LsSf https://astral.sh/uv/install.sh | sh
Setup
git clone https://github.com/semgrep/mythos-bench
cd mythos-bench
# Create .env with your OpenRouter key — never commit this file
echo 'OPENROUTER_API_KEY=sk-or-v1-...' > .env
For `harness.py`, sync the Python dependencies:
uv sync
For `cc_harness.go`, build the binary:
go build -o cc_harness cc_harness.go
Running
`cc_harness` (agentic, recommended)
# Dry run — print plan without invoking claude
./cc_harness -dry-run
# Full run, all models, all cases, 8 iterations each
./cc_harness
# Specific model and test case
./cc_harness -models anthropic/claude-opus-4-6 -cases openbsd-sack
# Whole-file mode with full repo context (slower, uses git clone)
./cc_harness -clone
# Reduce parallelism (default 10; lower if hitting rate limits)
./cc_harness -concurrency 3
# See all flags
./cc_harness -help
Key flags:
| Flag | Default | Des
Related Skills
ExoQuery Bench
Open benchmark for LLM assistants that turn plain-English astronomy questions into ADQL. Scores Claude, OpenAI
People Search Bench
The first open benchmark for evaluating AI-powered people search agents
10x Bench Kit
Create your own benchmark for coding AI Agents
Bench Watch
Launch or attach to a Plumbline benchmark slice, poll it to completion, and emit the canonical anti-Goodhart p
Klarion
AI-powered secret scanner for code and git history. Finds leaked API keys, tokens and credentials with near-ze
Spmind
[ICML 2026] Autonomous AI agent for end-to-end spatial proteomics analysis, with SP-Bench for agentic multiple
Related Agents
LLM AI Hunter
LLM and Agentic AI vulnerability specialist. Covers OWASP LLM Top 10 v2025 (LLM01-LLM10) and OWASP Agentic AI
Bench
API performance benchmarking — latency profiling, throughput testing, performance regression detection
Ecc Security Reviewer
Security vulnerability detection and remediation specialist. Use PROACTIVELY after writing code that handles u