Pharn Dev Eval — Development skill for Claude Code
Run a capability's eval LIVE via claude -p N times into isolated runs/, then COUNT structural pass/fail across the runs with the deterministic .dev/floor/check-variance.mjs.
How to install Pharn Dev Eval
Installs to ~/.claude/commands/pharn-dev-pharn-oss-pharn-dev-eval.md
mkdir -p ~/.claude/commands && curl -fsSL https://raw.githubusercontent.com/pharn-dev/pharn-oss/HEAD/.claude/commands/pharn-dev-eval.md -o ~/.claude/commands/pharn-dev-pharn-oss-pharn-dev-eval.md Restart Claude Code, or start a new session, for it to be picked up.
What Pharn Dev Eval does
description: "Run a capability's eval LIVE via claude -p N times into isolated runs/, then COUNT structural pass/fail across the runs with the deterministic .dev/floor/check-variance.mjs. The first live emission + the first variance measurement. flaky-structural = FAIL; semantic = advisory report." role: skill kind: pharn-owned trust: trusted model_tier: sonnet reads: [ "pharn/pharn-review/trust-fence/trust-fence.md", "pharn/pharn-review/trust-fence/evals/cases/case-injection-comme
Alternatives in Development
- Breach Check — HIBP k-anonymity check on a password wordlist 4.5k ★
- Agent Skill — A Claude Code plugin marketplace containing the ast-grep skill for powerful structural code search using Abstr 851 ★
- Skillcount — Skill Count — 技能数量统计与文档同步 835 ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Pharn Dev Grill
Interrogate an approved PLAN.md BEFORE /pharn-dev-build AND deterministically re-verify the plan's applied_les
Pharn Dev Plan
Plan ONE increment of PHARN. Discovery-first, grounded in live state, pins the architecture content-hash, halt
Warden Power
Zero-token power planner — from the agent's own recorded run-to-run variance, report the minimum detectable sa
Fable Eval
Run the fable-mode eval suite (probes → pairwise judge → report). Costs tokens — runs headless claude many tim
Kit Health
Run a self-assessment of the kit against its own philosophy. Checks file count, hook performance, source citat
Eval Prompts
Run a golden-set prompt eval suite against the pinned baseline; fail on any regression beyond threshold.
Related Agents
Flake Detector
Use when a test is suspected to be non-deterministic — passes sometimes, fails sometimes, or fails only in CI.
Robo Scout
Live web research for robotics. Verifies current versions, prices, lead times, part availability, new releases
Recorder
The review seat over consolidation — the pipeline's keystone, the one error class that COMPOUNDS. The engine's