Bench Capture — Development skill for Claude Code
Score the 2 bench variant folders for one size (build/acceptance/regression/conventions/capture + rubric + cross-solution review), record, report, and reset.
How to install Bench Capture
Installs to ~/.claude/commands/alfredoperez-speckit-companion-bench-capture.md
mkdir -p ~/.claude/commands && curl -fsSL https://raw.githubusercontent.com/alfredoperez/speckit-companion/HEAD/.claude/commands/bench-capture.md -o ~/.claude/commands/alfredoperez-speckit-companion-bench-capture.md Restart Claude Code, or start a new session, for it to be picked up.
What Bench Capture does
allowed-tools: Bash(node examples/todo-claude/bench/run-all.mjs:*), Agent, Workflow, AskUserQuestion description: Score the 2 bench variant folders for one size (build/acceptance/regression/conventions/capture + rubric + cross-solution review), record, report, and reset
Your task
After you've run a size through the two folders in VS Code, capture the evals, append to history, regenerate the report, and reset the folders for the next round.
1. Resolve the size
From `$ARGUMENTS`
Alternatives in Development
- Extraction Judge.Prompt — Extraction Judge — Fact Durability Rubric 189 ★
- Bench — Run performance benchmarks for RLM-Claude-Code 86 ★
- Batch Benchmark — Launch the 4-variant batch on the currently-free local GPUs (parallel when N>=4, otherwise sequential waves of 68 ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Bench Run All
Agent-driven bench round for one size — drive the 2 variant folders + judge + capture, hands-off
Bench Prep
Clean + arm the 2 bench variant folders for one size, ready to run in VS Code
Jev Gates
Three gates for any coding agent: a deterministic approval gate before irreversible actions, a completion gate
Score Opportunity
Calculate opportunity score using weighted rubric
Score Company Icp
Score company evidence {{company_evidence}} against this operator-authored rubric: {{icp_rubric}}. Apply its p
Score Persona Icp
Apply the operator-authored rubric {{persona_rubric}} to {{person_evidence}}. Sum only supported points, 0–100
Related Agents
Chief
Head of the PO Council and eight-seat design bench. Owns whether to convene, seat selection, order, named conf
Record Matcher
Scores candidate archival hits against a case identity. Name-variant aware, tolerant of date drift, weights fa
Bench
API performance benchmarking — latency profiling, throughput testing, performance regression detection