Since Cutoff Benchmark — Development skill for Claude Code
Pre-registered head-to-head benchmark of coding-agent context for post-cutoff library changes: Claude Code alone vs Context7 docs vs since-cutoff notes vs both (360 sessions, hidden tests, all transcr.
How to install Since Cutoff Benchmark
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open MohammadHijjawi97/since-cutoff-benchmark and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Since Cutoff Benchmark does
Pre-registered head-to-head benchmark of coding-agent context for post-cutoff library changes: Claude Code alone vs Context7 docs vs since-cutoff notes vs both (360 sessions, hidden tests, all transcripts).
Alternatives in Development
- Param Discover — Discover hidden HTTP parameters on a URL or list of URLs using Arjun (or x8 fallback) 4.5k ★
- Module Routes — Show all registered routes for a module 407 ★
- Obsidian Emerge — Surface unnamed patterns from your recent notes — recurring themes, hidden connections, and conclusions you ha 277 ★
README
since-cutoff head-to-head benchmark
When a Python project pins a library release that changed after a coding model's training cutoff, the model's memory of that API can be wrong. This repository holds the harness, the tasks, the frozen protocol and the complete records of a benchmark that measures, on the same tasks, how Claude Code does with:
- A: its usual tools only;
- B: plus the Context7 documentation MCP server;
- C: plus notes from since-cutoff in
CLAUDE.md; - D: plus both;
- E: plus Context7, with its use required by
CLAUDE.md(added after the pilot).
since-cutoff writes short notes about library APIs that changed after a model's training cutoff into `AGENTS.md` or `CLAUDE.md`, from a static diff of the package's public API between the release current at the cutoff and the pinned release. The benchmark was designed and run by the author of since-cutoff. The protocol, tasks, notes and analysis code were committed and tagged before the main run (an unsigned git tag; see [what was fixed in advance](#what-was-fixed-in-advance-and-what-was-added-later)), and every session's transcript is published here so each number can be checked.
A longer write-up is at .
Result of the main run (2026-09-28)
360 headless sessions: 24 tasks × 5 arms × 3 repetitions, Claude Code 2.1.283, model `claude-opus-5-5`, effort `medium`, on Windows, graded by hidden tests. Every planned run has a valid record (one harness error was re-run, see [deviations](#deviations-and-incidents)). The model's training cutoff was taken as 2026-06-30.
Numbers come from [`results/main-2026-09-28/report.json`](results/main-2026-09-28/report.json) (all tasks) and [`report.without-ant01-mcp01.json`](results/main-2026-09-28/report.without-ant01-mcp01.json) (the two pilot-exposed tasks removed). Ratios are geometric mea
Related Skills
Context7
Context7 — up-to-date library docs for your agent
08 Content Repurposing Engine
Turn one piece of content into a set of platform-native posts that each stand alone, all inheriting the same v
Pharn Dev Regress
Detect regressions OUTSIDE the just-built feature: re-run the existing deterministic suite (npm run check's ga
Blast Radius Bench
A benchmark for agentic coding-tool judgment under ambiguity: does the agent confirm before touching ambiguous
Cost Report
Pre-flight and post-flight cost visibility with dry-run, soft caps, and expensive-command warnings
Product Hunt Launch
Build a Product Hunt specific launch runbook, pre-launch prep, launch-day hour-by-hour plan, and post-launch f
Related Agents
Palantir
Researches a question on the open web (an advisory, a claim to fact-check, practice beyond the training cutoff
Websearcher
Targeted web lookup agent. Use proactively when the conversation requires external knowledge such as API refer
Content Author
Drafts one public piece — a launch post, release note, or documentation page — with every factual claim regist