Live Evaluator — Testing & QA agent for Claude Code
Phase 3 of the grounded-persona-eval method — drives one persona through a real, live interaction with the product under test and reports a candid, honestly-hedged verdict.
How to install Live Evaluator
Installs to ~/.claude/agents/qte77-agentic-grounded-persona-eval-live-evaluator.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/qte77/agentic-grounded-persona-eval/HEAD/.claude/agents/live-evaluator.md -o ~/.claude/agents/qte77-agentic-grounded-persona-eval-live-evaluator.md Restart Claude Code, or start a new session, for it to be picked up.
What Live Evaluator does
name: live-evaluator description: Phase 3 of the grounded-persona-eval method — drives one persona through a real, live interaction with the product under test and reports a candid, honestly-hedged verdict. One instance per persona, run in parallel.
live-evaluator
You run **Phase 3 (Live evaluation)** of the grounded-persona-eval method, for exactly **one** persona. If evaluating multiple personas, launch one instance of this agent per persona, in parallel — never in series, and neve
Alternatives in Testing & QA
- Logic Hacker — Red Team persona for Business Logic and Auth manipulation 548 ★
- Witness — Independently witness that an Allium loop's convergence claim is true and was reached honestly 474 ★
- Writ Test Writer — Writes test skeleton files with method signatures and assertions based on an approved plan 187 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
MCP Dogfooder
Drives the shipped MCP debugger tools as a real user would — launch, breakpoints, stepping, locals, object exp
Manual QA
Use when a change needs to be exercised in a real running app — not unit tests, but a human-style check of the
Worca Cc Manual Web UI Testing
Manual web UI testing agent for the orchestrator pipeline. Drives the RUNNING web UI through the manual test c
Fixme Review Code
Reviews code produced by plan execution. Finds real bugs, gaps, test issues, and inconsistencies. Read-only -
Stx Reviewer
Multi-agent wave Reviewer persona. Reads the Dev's diff plus the task spec from architecture-verse.html and em
Kavach Confirm Reporter
KAVACH live-validation reporting agent. Aggregates every confirm_status verdict kavach-poc-executor and kavach
Related Skills
Skill Harness
I wanted to know whether a skill was actually any good. This runs the same task with the skill and without it,
Agentic Eval
Niche-agnostic agentic evaluator using CLEAR v2.0 framework — 6-domain assessment, 8 analysis dimensions, 6-ti
Claude Interaction Tests
Test whether your CLAUDE.md actually works. A Claude Code skill that restructures agent-facing docs (CLAUDE.md