stas4000

Decisionbench — Testing skill for Claude Code

Testing community

One public test, any typed-decision model: accuracy, latency, cost, and accuracy when it is sure.

How to install Decisionbench

This entry records only its repository, not the path inside it, so there is no exact command to give. Open stas4000/decisionbench and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Decisionbench does

One public test, any typed-decision model: accuracy, latency, cost, and accuracy when it is sure. Jev vs Claude Opus 5.5 on Banking77.

Alternatives in Testing

  • Mda — MDA Open Spec — a Markdown superset for agent-facing documents 615 ★
  • Next Evals OSS — Evals for Next.js up to 15.5.6 to test AI model competency at Next.js 310 ★
  • GAAI Framework — 95+ Drop-in governance layer (.gaai/ folder) -- backlog authorization, cross-session memory, decision tracking 113 ★

README

decisionbench

One public labeled test, any typed-decision model, the same four numbers for all of them: accuracy, latency (p50/p95), cost per 1,000 decisions, and accuracy when the model says it is sure.

New "typed decision" models keep shipping with speed numbers (yes/no in 80 ms, a choice in 100 ms). Speed is half the story. This runs them on the same public cases so the other half is on the table too.

Built with Claude Opus 5.5.

![preview](docs/preview.png)

Result, one run, 27 Sept 2026

154 customer messages from the public Banking77 test set, 2 per intent, 77 possible intents, seed 7. Real API calls, one run each.

model right cost for 154 per 1,000 p50 p95
TypeSafe Jev (jev-latest) 79.9% $0.011 $0.07 0.26 s 0.33 s
Claude Opus 5.5 (Claude Code CLI) 83.8% $2.11 $13.68 4.17 s 9.40 s
route: Jev when sure (>= 0.8), else Opus 5.5 83.1% $0.51 $3.33
  • Jev said it was at least 80% sure on 122 of 154 messages (79%), and was right on 88.5% of those, which is above Opus 5.5 overall.
  • Routing only the 32 unsure messages to Opus 5.5 gets within 0.7 points of Opus alone for a quarter of the cost.
  • Opus cost is what the CLI reported, including its own prompt overhead. Jev cost is input tokens at $0.042 per million.

Raw per-message results are in `runs/banking77-154/`.

Run it

python3 decisionbench.py --models jev,claude --n 154

Stdlib only, Python 3.10+. The test set downloads on first run.

adapter needs
jev TYPESAFE_API_KEY
claude the claude CLI on PATH, CLAUDE_MODEL (default claude-opus-5-5)
openai any OpenAI-compatible endpoint: OPENAI_BASE_URL, OPENAI_MODEL, optional OPENAI_API_KEY. Use it for a local model served by vLLM, llama.cpp or Ollama

`--sure 0.8` sets the confidence line, `--ipv4` for hosts with a broken IPv6 route. A failed call counts as a miss, never dropped. Adding a model is one function that