Sdlc Evals — Testing skill for Claude Code
/sdlc-evals — Author the Versioned Golden Set for an LLM-Powered Spec.
How to install Sdlc Evals
Installs to ~/.claude/skills/mckruz-claude-code-sdlc-sdlc-evals/SKILL.md
mkdir -p ~/.claude/skills/mckruz-claude-code-sdlc-sdlc-evals && curl -fsSL https://raw.githubusercontent.com/MCKRUZ/claude-code-sdlc/HEAD/commands/sdlc-evals.md -o ~/.claude/skills/mckruz-claude-code-sdlc-sdlc-evals/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
What Sdlc Evals does
/sdlc-evals — Author the Versioned Golden Set for an LLM-Powered Spec
Author the **golden set** — the acceptance criteria for probabilistic behavior — for a spec on an `llm_powered` channel. Bizreq owns the scenarios, Data owns the data; the two are authored collaboratively. This command is a thin **wrapper around the existing `eval-builder` harness skill**: it adds **no new script**, authors **no CI YAML**, and writes a versioned `golden-set.yaml` next to the spec. Interview-driven like `/sd
Alternatives in Testing
- Notion Spec To Implementation — Convert Notion specs into linked implementation plans and tasks 14.6k ★
- Claude Code Spec Workflow — Automated Kiro-style Spec workflow for Claude Code 3.6k ★
- Cc Sdd — Spec-driven development (SDD) for your team's workflow 3.1k ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Readme Generator
Generate professional README.md with 16:9 infographics, SEO-optimized metadata, and structured author sections
Pharn Spec
Turn a user's prose intent into a structured, human-approved features/ /SPEC.md — the head of the product pipe
Fia Harness
Local, dependency-free spec-driven development harness for AI agents: spec-first, machine-validated state and
Gaia Spec
Author an immutable SPEC artifact through Socratic discovery (spec-kit wrapper), then STOP. Terminal, never ru
Spec Dc
Author a behavior specification (the WHAT) — step 1 of spec → plan → execute.
Marketing Design
Direct invocation of Marketing Designer. Author or revise a marketing-design-spec for a project's public marke
Related Agents
Experiment Reporter
Background experiment report consolidator — given a canonical experiment_id and disk pointers to runs, arms, e
Eval Author
Authoring role for the D365FO agent eval loop catalog. Drafts a new eval/cases/ .json spec (valid against eval
AI Eval Designer
Use this agent to design a risk-tiered evaluation set for an AI feature. Trigger when the user says "design ev