AI Agent Eval Writer — Testing & QA agent for Claude Code
Use this agent when you need to write evaluation tests for AI agents using Evalite or Autoevals frameworks.
How to install AI Agent Eval Writer
Installs to ~/.claude/agents/januarylabs-deepagents-ai-agent-eval-writer.md
mkdir -p ~/.claude/agents && curl -fsSL https://raw.githubusercontent.com/JanuaryLabs/deepagents/HEAD/.claude/agents/ai-agent-eval-writer.md -o ~/.claude/agents/januarylabs-deepagents-ai-agent-eval-writer.md Restart Claude Code, or start a new session, for it to be picked up.
What AI Agent Eval Writer does
name: ai-agent-eval-writer description: Use this agent when you need to write evaluation tests for AI agents using Evalite or Autoevals frameworks. This includes creating evaluation suites to test LLM outputs, agent behaviors, response quality, factual accuracy, or any other AI system performance metrics. Examples of when to invoke this agent:\n\n\nContext: The user has just finished implementing an AI agent or LLM-based feature and wants to ensure it performs correctly.\nuser: "I j
Alternatives in Testing & QA
- Rivetkit Driver Test Writer — Use this agent when the user needs to write, modify, or debug driver tests for RivetKit 6.1k ★
- Testing Expert — Use this agent for Output.ai testing strategies including Vitest configuration, Temporal workflow testing, LLM 434 ★
- Module Test Writer — Use this agent to write feature tests for LaraDashboard module CRUD operations, services, policies, and Livewi 407 ★
Full documentation available on GitHub
View Source RepositoryRelated Agents
Logic Architect
Extract capabilities and behaviors from spec. Outputs JSON that MUST validate against schemas/capability-map.s
Go Test Writer Assistant
Write comprehensive Go tests using Ginkgo/Gomega following project patterns. Generates test suites, uses Count
AI Prompt Architect
Designs and versions LLM system prompts for ai-system / agent-product archetypes. Outputs docs/decisions/ADR-{
Eval Auditor
Audits an evaluation setup (benchmark, A/B test, or model comparison) for methodology errors that would invali
AI Eval Designer
Use this agent to design a risk-tiered evaluation set for an AI feature. Trigger when the user says "design ev
Brain Eval Engineer
Evaluation engineer — question set tooling, layered metrics (harvest/graph/retrieval/answer), fixed-strategy a
Related Skills
E2e Test Writer
Write end-to-end tests using Playwright or Cypress
Ashlr Eco Mode
Toggle eco mode — aggressive token-saving behaviors that trade some response richness for lower session cost.
Coverage Gaps
Find production behaviors that have no evaluation measuring them, and prioritize which judges to build.