Benchmark Rerankers — Development skill for Claude Code
Measure whether adding a reranker actually improves retrieval, by scoring reranked vs. un-reranked results on a labeled query set.
How to install Benchmark Rerankers
Installs to ~/.claude/skills/imtiazrayhan-agentscamp-library-benchmark-rerankers/SKILL.md
mkdir -p ~/.claude/skills/imtiazrayhan-agentscamp-library-benchmark-rerankers && curl -fsSL https://raw.githubusercontent.com/imtiazrayhan/agentscamp-library/HEAD/commands/benchmark-rerankers.md -o ~/.claude/skills/imtiazrayhan-agentscamp-library-benchmark-rerankers/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
What Benchmark Rerankers does
description: "Measure whether adding a reranker actually improves retrieval, by scoring reranked vs. un-reranked results on a labeled query set." argument-hint: "" allowed-tools: "Read, Grep, Glob, Bash" model: sonnet
Scope
Treat `$ARGUMENTS` as the retrieval setup to benchmark — a path to an eval set and retrieval results, or a description of the pipeline (retriever, candidate count, reranker). Restate what you
Alternatives in Development
- Skill Creator — Create new skills, modify existing skills, and measure skill performance 94.1k ★
- Fidelity — Measure how faithfully a clone reproduces a site — pixel-diff plus motion-fidelity into one 0-100 score, a let 3.6k ★
- Evo — turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then run 1.4k ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
Prompt Injection Benchmark
A reproducible prompt-injection benchmark that measures which defenses actually work: each payload is replayed
Long Context Retrieval
Read large files and broad search results progressively instead of blowing up the live context window.
MCP Recall
mcp-recall compresses MCP tool outputs (94 KB → 3.5 KB · 96%) and stores full results in SQLite for retrieval
Forge Benchmark
Measure validation posture across 5 dimensions with trend tracking
Auxi Rn Patterns
React Native + TanStack Query patterns specific to the Auxi mobile app. Use when adding screens, services, or
Contributing To Smart Ralph
"I'm learnding!" - You, after reading this guide First off, thanks for wanting to contribute! This project wel
Related Agents
Ndv Signal
Metrics skeptic. Use when reviewing engineering KPIs, OKRs, sprint velocity, test coverage targets, DORA metri
Selfish Reputation
Measure whether the body of work is increasing qualified trust, recognition, referrals, repeat engagement, and
Eval Auditor
Audits an evaluation setup (benchmark, A/B test, or model comparison) for methodology errors that would invali