Decisionbench — Testing skill for Claude Code
One public test, any typed-decision model: accuracy, latency, cost, and accuracy when it is sure.
How to install Decisionbench
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open stas4000/decisionbench and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Decisionbench does
One public test, any typed-decision model: accuracy, latency, cost, and accuracy when it is sure. Jev vs Claude Opus 5.5 on Banking77.
Alternatives in Testing
- Mda — MDA Open Spec — a Markdown superset for agent-facing documents 615 ★
- Next Evals OSS — Evals for Next.js up to 15.5.6 to test AI model competency at Next.js 310 ★
- GAAI Framework — 95+ Drop-in governance layer (.gaai/ folder) -- backlog authorization, cross-session memory, decision tracking 113 ★
README
decisionbench
One public labeled test, any typed-decision model, the same four numbers for all of them: accuracy, latency (p50/p95), cost per 1,000 decisions, and accuracy when the model says it is sure.
New "typed decision" models keep shipping with speed numbers (yes/no in 80 ms, a choice in 100 ms). Speed is half the story. This runs them on the same public cases so the other half is on the table too.
Built with Claude Opus 5.5.

Result, one run, 27 Sept 2026
154 customer messages from the public Banking77 test set, 2 per intent, 77 possible intents, seed 7. Real API calls, one run each.
| model | right | cost for 154 | per 1,000 | p50 | p95 |
|---|---|---|---|---|---|
TypeSafe Jev (jev-latest) |
79.9% | $0.011 | $0.07 | 0.26 s | 0.33 s |
| Claude Opus 5.5 (Claude Code CLI) | 83.8% | $2.11 | $13.68 | 4.17 s | 9.40 s |
| route: Jev when sure (>= 0.8), else Opus 5.5 | 83.1% | $0.51 | $3.33 |
- Jev said it was at least 80% sure on 122 of 154 messages (79%), and was right on 88.5% of those, which is above Opus 5.5 overall.
- Routing only the 32 unsure messages to Opus 5.5 gets within 0.7 points of Opus alone for a quarter of the cost.
- Opus cost is what the CLI reported, including its own prompt overhead. Jev cost is input tokens at $0.042 per million.
Raw per-message results are in `runs/banking77-154/`.
Run it
python3 decisionbench.py --models jev,claude --n 154
Stdlib only, Python 3.10+. The test set downloads on first run.
| adapter | needs |
|---|---|
jev |
TYPESAFE_API_KEY |
claude |
the claude CLI on PATH, CLAUDE_MODEL (default claude-opus-5-5) |
openai |
any OpenAI-compatible endpoint: OPENAI_BASE_URL, OPENAI_MODEL, optional OPENAI_API_KEY. Use it for a local model served by vLLM, llama.cpp or Ollama |
`--sure 0.8` sets the confidence line, `--ipv4` for hosts with a broken IPv6 route. A failed call counts as a miss, never dropped. Adding a model is one function that
Related Skills
Typesafe Jev
Field notes, runnable scripts and an agent skill for TypeSafe Jev, the System One decision model. Measured eva
Jev Cold Start Prior
Can a TypeSafe Jev prior read from a README on day one predict which new agent-skill repos gain stars? Zero-sh
REST API
Generate a modern ACS v1 Public REST API resource (annotation-based @EntityResource / @RelationshipResource wi
PM Working Backwards Agent
Multi-agent CrewAI pipeline that turns a product problem statement into a research brief, PRFAQ, BRD, and buil
Autopus Adk
Autopus-ADK is of the agents, by the agents. for the agents. Multi-model orchestration (consensus/pipeline/deb
Test Example
When we make significant changes to the codebase, we want to make sure everything is running as intended. To d
Related Agents
Brain Eval Engineer
Evaluation engineer — question set tooling, layered metrics (harvest/graph/retrieval/answer), fixed-strategy a
Effort Low
Pins LOW reasoning effort; pair with a per-call model param. For mechanical, single-rule, high-volume fan-out
Documentation Manager
Expert documentation specialist. Proactively updates documentation when code changes are made, ensures README