MohammadHijjawi97

Since Cutoff Benchmark — Development skill for Claude Code

Development community

Pre-registered head-to-head benchmark of coding-agent context for post-cutoff library changes: Claude Code alone vs Context7 docs vs since-cutoff notes vs both (360 sessions, hidden tests, all transcr.

How to install Since Cutoff Benchmark

This entry records only its repository, not the path inside it, so there is no exact command to give. Open MohammadHijjawi97/since-cutoff-benchmark and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Since Cutoff Benchmark does

Pre-registered head-to-head benchmark of coding-agent context for post-cutoff library changes: Claude Code alone vs Context7 docs vs since-cutoff notes vs both (360 sessions, hidden tests, all transcripts).

Alternatives in Development

  • Param Discover — Discover hidden HTTP parameters on a URL or list of URLs using Arjun (or x8 fallback) 4.5k ★
  • Module Routes — Show all registered routes for a module 407 ★
  • Obsidian Emerge — Surface unnamed patterns from your recent notes — recurring themes, hidden connections, and conclusions you ha 277 ★

README

since-cutoff head-to-head benchmark

When a Python project pins a library release that changed after a coding model's training cutoff, the model's memory of that API can be wrong. This repository holds the harness, the tasks, the frozen protocol and the complete records of a benchmark that measures, on the same tasks, how Claude Code does with:

  • A: its usual tools only;
  • B: plus the Context7 documentation MCP server;
  • C: plus notes from since-cutoff in CLAUDE.md;
  • D: plus both;
  • E: plus Context7, with its use required by CLAUDE.md (added after the pilot).

since-cutoff writes short notes about library APIs that changed after a model's training cutoff into `AGENTS.md` or `CLAUDE.md`, from a static diff of the package's public API between the release current at the cutoff and the pinned release. The benchmark was designed and run by the author of since-cutoff. The protocol, tasks, notes and analysis code were committed and tagged before the main run (an unsigned git tag; see [what was fixed in advance](#what-was-fixed-in-advance-and-what-was-added-later)), and every session's transcript is published here so each number can be checked.

A longer write-up is at .

Result of the main run (2026-09-28)

360 headless sessions: 24 tasks × 5 arms × 3 repetitions, Claude Code 2.1.283, model `claude-opus-5-5`, effort `medium`, on Windows, graded by hidden tests. Every planned run has a valid record (one harness error was re-run, see [deviations](#deviations-and-incidents)). The model's training cutoff was taken as 2026-06-30.

Numbers come from [`results/main-2026-09-28/report.json`](results/main-2026-09-28/report.json) (all tasks) and [`report.without-ant01-mcp01.json`](results/main-2026-09-28/report.without-ant01-mcp01.json) (the two pilot-exposed tasks removed). Ratios are geometric mea