Eval From Trace
Description
--- description: Build an LLM-as-a-Judge eval grounded in real traces from the Progress Observability Platform. argument-hint: [application/service name and optional symptom, e.g. "checkout-agent, wrong tool calls"] --- Use the `generate-eval` skill, Workflow A (from traces). Target: $ARGUMENTS Steps: 1. Survey the system's recent traffic with the metadata-only observability tools (`list_observations`) over the last 24–72h. 2. Choose a single failure mode and config from what the traces actua
Installation
Installs to ~/.claude/skills/observability-oss-progress-observability-plugin-eval-from-trace/SKILL.md
mkdir -p ~/.claude/skills/observability-oss-progress-observability-plugin-eval-from-trace && curl -fsSL https://raw.githubusercontent.com/observability-oss/progress-observability-plugin/HEAD/commands/eval-from-trace.md -o ~/.claude/skills/observability-oss-progress-observability-plugin-eval-from-trace/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
Full documentation available on GitHub
View Source RepositoryRelated Skills
Agency Agents
A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy inject
AI Awesome Llm Apps
100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
AI Firecrawl
🔥 The API to search, scrape, and interact with the web for AI
AI Artifacts Builder
Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech
AI Headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agen
AI CrewAI
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewA
AI