OSS Migration Eval — Data skill for Claude Code
Decide whether to switch an LLM pipeline to a cheaper/open-source model — honest cost-vs-accuracy verdict with error bars.
How to install OSS Migration Eval
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open BayramAnnakov/oss-migration-eval and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What OSS Migration Eval does
Decide whether to switch an LLM pipeline to a cheaper/open-source model — honest cost-vs-accuracy verdict with error bars. A Claude Code skill.
Alternatives in Data
- Token Dashboard — See where Claude Code is burning tokens - turn raw JSONL transcripts into local cost analytics, hotspot views 669 ★
- Dbt Migration — Migrate SQL projects to dbt 276 ★
- Tokdash — Agent Dashboard: Visualization and analytics for Sessions and Quota Usage 70 ★
README
oss-migration-eval
**Decide — rigorously — whether to move an expensive LLM pipeline to a cheaper or open-source model.** A Claude Code skill that estimates *cost savings vs accuracy loss*, finds the most probable replacement models, and produces an honest go/no-go verdict with error bars.
The uncomfortable truth this skill is built around: **most eval "winners" are single-sample noise.** A real experiment crowned a decisive 9.00/10 winner off one generation per cell; a 5-seed rigor pass dropped it to 7.00 ± 2.0 — mid-pack in a six-way statistical tie. Within-model variance across seeds nearly equalled the entire between-model ranking. So the honest output is often *"no clear winner — here's the tradeoff, and here are the models you can safely rule out."* Saying that is a success, not a failure.
What it does
Given a recurring, expensive LLM task (lead scoring, extraction, classification, outbound copy, …), it:
- Finds the fattest tasks worth migrating (pairs with `token-audit`).
- Frames the accuracy metric — deterministic for structured output, a blind panel for subjective.
- Builds a human-validated golden set on your frontier model, including hard and false-trigger trap cases.
- Shortlists candidates live from quality (Artificial Analysis) and usage (OpenRouter) leaderboards — never a hardcoded list.
- Sweeps each candidate provider-pinned across multiple seeds, recording measured cost + latency.
- Grades with error bars and reports statistically tied vs clearly worse — you can only reliably name the losers, not a winner.
- Optionally optimizes the gap with a prompt-tuning loop (train/test/val discipline; pairs with
autoresearch). - Deploys via orchestration — the smart model orchestrates, the cheap model does the narrow task; local via Ollama; fine-tuning as a last resort.
Why it's different — the honesty guardrails
Baked in, because a naive eval
Related Skills
Fixer Plan
You are triaging a single pull-request review comment for an automated pipeline. Your ONLY job is to decide wh
Roblox Ideation Skill
Game ideation & concept-design companion for Roblox — a Claude Code skill. The front of the studio pipeline: d
Cost Of Remembering
The Cost of Remembering: filesystem memory matches long-context accuracy on LongMemEval while reading 97% fewe
Review Verdict
SkillForge pipeline.md 流派 review verdict 模板(PASS / blocker / warning / nit + Stage 1/2 + r2+ prior items verif
WhereMyTokens
new Windows system tray app for monitoring Claude Code token usage in real time. Per-session token counts, cos
Borg
One brain, many hands: shared local memory (mem0 + Graphiti) for Claude/Codex/Grok agents, a grammar-shim data
Related Agents
Network Config Reviewer
Network configuration security and correctness auditor. Invoke when given a router or switch configuration to
OSS
Open-source maintainer of redbar — repository governance, not the prose (that's the scribe's job). Invoke for
OSS Maintainer
Mantiene la higiene open source de las piezas del paraguas ai-native (cli, skills, mcp). Úsalo antes de public