timerise-ai

Skill Eval Loop — Development skill for Claude Code

Development community

Agent Skill: harden another skill from its automatic agent evals.

How to install Skill Eval Loop

This entry records only its repository, not the path inside it, so there is no exact command to give. Open timerise-ai/skill-eval-loop and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Skill Eval Loop does

Agent Skill: harden another skill from its automatic agent evals. Score every Claude Code, Codex and Gemini CLI run against a fixed fidelity rubric read from the agent's log, fix the skill at the root cause, cut a patch release, and loop until every agent scores full

Alternatives in Development

  • Cs2 Knife — Create a high-fidelity procedural Three.js 3D reconstruction of 14.1k ★
  • Fidelity — Measure how faithfully a clone reproduces a site — pixel-diff plus motion-fidelity into one 0-100 score, a let 3.6k ★
  • Quick Eval — Quick job evaluation 473 ★

README

skill-eval-loop

[![Agent Skills](https://img.shields.io/badge/Agent_Skills-open_format-059669)](https://agentskills.io) [![skills.sh](https://img.shields.io/badge/skills.sh-npx_skills_add-059669)](https://www.skills.sh) [![Claude Code](https://img.shields.io/badge/Claude_Code-compatible-059669)](https://docs.claude.com/en/docs/claude-code/skills) [![Codex CLI](https://img.shields.io/badge/Codex_CLI-compatible-059669)](https://developers.openai.com/codex/skills) [![Gemini CLI](https://img.shields.io/badge/Gemini_CLI-compatible-059669)](https://github.com/google-gemini/gemini-cli/blob/main/docs/cli/skills.md)

An [Agent Skill](https://agentskills.io) for skill maintainers. It hardens another skill, one that builds a module for **Next.js App Router** apps, from its automatic agent evals: score every run against a fixed fidelity rubric read from the agent's own log, fix the skill at the root cause of each deviation, cut a patch release, let the release re-run the evals, and repeat until Claude Code, Codex and Gemini CLI all score full.

**An eval that passes its checks has not proved the agent used the skill as written.** The deviations live in the agent's log, not in the result's checks, and each one traces back to a sentence in the skill that allowed it. Score the log, fix the sentence, and let the next release's eval prove the fix.

This skill was written by the maintainer who has run the loop on a published skill, from its first scored release to a unanimous full score. What it carries holds by construction: a rubric fixed before round one and the same for every agent, so scores compare across rounds; every agent's template edit reproduced against the template before it is adopted; a template check that type-checks the templates and matches the documented test count before any round ships; and a stop rule that is set in advance and never moved. [`references/provenance.md`](references/provenance.md) has the record.

Install

One command, via the [skills.sh](htt