Skill Eval Loop — Development skill for Claude Code
Agent Skill: harden another skill from its automatic agent evals.
How to install Skill Eval Loop
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open timerise-ai/skill-eval-loop and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Skill Eval Loop does
Agent Skill: harden another skill from its automatic agent evals. Score every Claude Code, Codex and Gemini CLI run against a fixed fidelity rubric read from the agent's log, fix the skill at the root cause, cut a patch release, and loop until every agent scores full
Alternatives in Development
- Cs2 Knife — Create a high-fidelity procedural Three.js 3D reconstruction of 14.1k ★
- Fidelity — Measure how faithfully a clone reproduces a site — pixel-diff plus motion-fidelity into one 0-100 score, a let 3.6k ★
- Quick Eval — Quick job evaluation 473 ★
README
skill-eval-loop
[](https://agentskills.io) [](https://www.skills.sh) [](https://docs.claude.com/en/docs/claude-code/skills) [](https://developers.openai.com/codex/skills) [](https://github.com/google-gemini/gemini-cli/blob/main/docs/cli/skills.md)
An [Agent Skill](https://agentskills.io) for skill maintainers. It hardens another skill, one that builds a module for **Next.js App Router** apps, from its automatic agent evals: score every run against a fixed fidelity rubric read from the agent's own log, fix the skill at the root cause of each deviation, cut a patch release, let the release re-run the evals, and repeat until Claude Code, Codex and Gemini CLI all score full.
**An eval that passes its checks has not proved the agent used the skill as written.** The deviations live in the agent's log, not in the result's checks, and each one traces back to a sentence in the skill that allowed it. Score the log, fix the sentence, and let the next release's eval prove the fix.
This skill was written by the maintainer who has run the loop on a published skill, from its first scored release to a unanimous full score. What it carries holds by construction: a rubric fixed before round one and the same for every agent, so scores compare across rounds; every agent's template edit reproduced against the template before it is adopted; a template check that type-checks the templates and matches the documented test count before any round ships; and a stop rule that is set in advance and never moved. [`references/provenance.md`](references/provenance.md) has the record.
Install
One command, via the [skills.sh](htt
Related Skills
Agent Evals Playground
A shopping agent, a labelled eval suite, and the wiring to score it against a real cluster instead of a fixtur
Harden Containers
Pin base images by digest, enforce non-root, and harden Dockerfiles
Eval Advisory
Eval advisory is a skill for planning, reviewing, and developing your evals.
Raven Harden
Reviews security_log.md and promotes accumulated observations into
Gallery
Build a static, shareable gallery from your fidelity reports — an index of score cards plus a permalink page (
Bench Capture
Score the 2 bench variant folders for one size (build/acceptance/regression/conventions/capture + rubric + cro
Related Agents
Mutation Hardener
Dual-mode mutation-testing agent. AUDIT mode runs the mutation tool over a scope and reports surviving mutants
AI Eval Designer
Use this agent to design a risk-tiered evaluation set for an AI feature. Trigger when the user says "design ev
Eval Failure Analyzer
Analyze Logic-Lens benchmark/eval failures. Use after running content-evals, or when pointed at a skills-works