Run Benchmarks — Development skill for Claude Code
Run Benchmarks — GenAIIDP empirical config/scaling suite.
How to install Run Benchmarks
Installs to ~/.claude/skills/aws-solutions-library-samples-accelerated-intelligent-document-processing-on-aws-run-benchmarks/SKILL.md
mkdir -p ~/.claude/skills/aws-solutions-library-samples-accelerated-intelligent-document-processing-on-aws-run-benchmarks && curl -fsSL https://raw.githubusercontent.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/HEAD/.claude/skills/run-benchmarks.md -o ~/.claude/skills/aws-solutions-library-samples-accelerated-intelligent-document-processing-on-aws-run-benchmarks/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
What Run Benchmarks does
Run Benchmarks — GenAIIDP empirical config/scaling suite
Use this skill to run the benchmark suite in `benchmarks/` — an end-to-end, ground-truth study across document types/sizes × configuration options that quantifies success/completeness/accuracy/calibration/latency/tokens/cost. Use it to (a) regenerate the results paper for a release, or (b) gate a code change by comparing against the committed baseline.
Read `benchmarks/matrices/METHODOLOGY.md` first — it defines the matrices, scori
Alternatives in Development
- Engage.Crash — Crash → root cause → reachability → empirical exploitability verdict (native bugs) 348 ★
- Scaling Prompts — Scaling Prompts — Road to 1M+ Users & Beyond CoinGecko 313 ★
- Contract Review Skill — Review legal contracts for risks, extract key terms, and suggest redlines 192 ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Show Config
Display current dev-suite configuration with enabled components
Jmunch Console
Local, opt-in control-plane GUI for the jMunch suite (config, index/watcher health, savings/ROI, sessions, age
Flow Elaboration To Construction
Orchestrate Elaboration→Construction phase transition with iteration planning, team scaling, and full-scale de
Lightning Channel Factories
Technical reference on Lightning Network channel factories, multi-party channels, LSP architectures, and Bitco
Lightning Architecture Review
Review Bitcoin Lightning Network protocol designs, compare channel factory approaches, and analyze Layer 2 sca
02a Empirical Reality Check
Stage 2a: Empirical Reality Check / Institutional Context Check
Related Agents
Certora Mutation Testing
Takes a Certora configuration and invariant suite, generates mutation campaigns correctly with certoraMutate a
Performance Benchmarker
Use for designing and running performance benchmarks with actionable reporting.
Eval Failure Analyzer
Analyze Logic-Lens benchmark/eval failures. Use after running content-evals, or when pointed at a skills-works