Letta Evals — Testing skill for Claude Code
Evaluation kit for testing stateful agents.
How to install Letta Evals
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open letta-ai/letta-evals and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Letta Evals does
Evaluation kit for testing stateful agents.
Alternatives in Testing
- Webapp Testing — Test local web applications using Playwright for UI verification and debugging 94.1k ★
- Fix Issue — by metabase - Addresses GitHub issues by taking issue number as parameter, analyzing context, implementing sol 46.5k ★
- Playwright — claude-plugins-official Browser automation, E2E testing, screenshots 29.4k ★
README
Letta Evals
Letta Evals is a framework for evaluating [Letta](https://github.com/letta-ai/letta) and Letta Code agents. It lets you define an evaluation suite with a dataset, target, extractors, graders, and a reward contract, then run that suite against one or more model configurations.
If you are building agentic systems, high-quality evals are one of the fastest ways to understand how model versions, prompts, tools, or agent configuration changes affect your product.
Requirements
- Python 3.11+
- A running Letta server, either:
- Self-hosted: follow the Letta installation guide, or
- Letta Cloud: create an account at app.letta.com and set:
Then useexport LETTA_API_KEY=your-api-key export LETTA_PROJECT_ID=your-project-idbase_url: https://api.letta.com/in your suite YAML, or pass--base-url https://api.letta.com/on the CLI.
- Provider API keys for the models you use, such as
OPENAI_API_KEY,ANTHROPIC_API_KEY, orGOOGLE_API_KEY.
Installation
For local development or custom eval authoring, clone this repository and install with dev dependencies:
uv sync --extra dev
To run existing evals without editing the repo:
pip install letta-evals
Quick start
- Create a dataset (
dataset.jsonl):
{"input": "What's the capital of France?", "ground_truth": "Paris"}
{"input": "Calculate 2+2", "ground_truth": "4"}
- Create a suite (
suite.yaml):
name: my-eval-suite
dataset: dataset.jsonl
target:
kind: letta_code
model_handles:
- openai/gpt-4.1-mini
base_url: http://localhost:8283
graders:
correctness:
kind: tool
function: contains
extractor: la
Related Skills
Next.js Evals Skills
Next.js evaluation and testing skills
Evaluation Framework
Test and validate changes to the resource kit systematically. Activate when testing new features, reviewing PR
Validate Hypothesis
Business hypothesis validation skill. Validates ideas through 6 phases (origin check, market confirmation, int
Agent Eval Design
Design an evaluation for an AI agent or LLM feature: what to test, how to grade it, and how to catch regressio
AI Qe Agent
AI QE Agent with LLM Evaluation Layer — catches hallucinations, monitors chain consistency, self-heals Playwri
Eval Suite
Re-score tiny-spec against the SDD evaluation rubric, record it, and report the delta vs the last run.
Related Agents
Browser Testing
Browser automation and testing specialist. Uses playwright-cli (stateful Bash CLI) for deterministic scripted
ML Engineer
Use for ML/AI model work — training, fine-tuning, evaluation, RAG, agents, embeddings, evals, deployment, MLOp
Statistician
Statistical rigor advisor: experimental design, signal-vs-noise, pressure-testing metric claims, sample-size a