Sw Eval — Development skill for Claude Code
Run Specwright eval suite.
How to install Sw Eval
Installs to ~/.claude/commands/obsidian-owl-specwright-sw-eval.md
mkdir -p ~/.claude/commands && curl -fsSL https://raw.githubusercontent.com/Obsidian-Owl/specwright/HEAD/.claude/commands/sw-eval.md -o ~/.claude/commands/obsidian-owl-specwright-sw-eval.md Restart Claude Code, or start a new session, for it to be picked up.
What Sw Eval does
Run Specwright eval suite.
Alternatives in Development
- Run Benchmarks — Run Benchmarks — GenAIIDP empirical config/scaling suite 295 ★
- Opencode — A powerful, custom opencode configuration, complete with a suite of agents, commands, rules, skills, and a pre 133 ★
- A B Copy — Use when you want a marketing-copy A/B — spawns two writer sibling sessions via CCC that draft the SAME messag 130 ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Agent Evals Playground
A shopping agent, a labelled eval suite, and the wiring to score it against a real cluster instead of a fixtur
Spar
Make the models argue before you build — Codex attacks the plan in bounded read-only rounds, then one builds a
Warden Bench
Run the golden benchmark suite for one token-warden agent (or all) and compare results against the frozen run1
Fable Eval
Run the fable-mode eval suite (probes → pairwise judge → report). Costs tokens — runs headless claude many tim
Eval Prompts
Run a golden-set prompt eval suite against the pinned baseline; fail on any regression beyond threshold.
Pair Verify
Use after you fix a bug and before you claim it fixed — it spawns ONE skeptic sibling session via CCC that mus
Related Agents
Bench Reviewer
Reviews eval suite quality — fixture/rubric consistency, scoring accuracy, false positive traps, difficulty ca
Claude Grader
Grades one extension test run against its assertions and writes grading.json. Use after an eval run finishes,
Lead Discovery
Discovery pipeline lead. Spawns scouts to investigate issue backlog, corpus gaps, and feature gaps. Promotes s