Tb2 Sandwich Recency — Development skill for Claude Code
Terminal-Bench 2.0 submission: Sandwich+Recency Ensemble agent achieving 99.55% (443/445 trials) with Claude Sonnet 4.5. Based on Harbor framework (laude-institute).
How to install Tb2 Sandwich Recency
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open K-sushi/tb2-sandwich-recency and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Tb2 Sandwich Recency does
Terminal-Bench 2.0 submission: Sandwich+Recency Ensemble agent achieving 99.55% (443/445 trials) with Claude Sonnet 4.5. Based on Harbor framework (laude-institute).
Alternatives in Development
- GenericAgent — Self-evolving agent: grows skill tree from 3.3K-line seed, achieving full system control with 6x less token co 6.1k ★
- /report — Generate a submission-ready bug bounty report 1.8k ★
- Axiom Marketplace Submission Guide — Axiom Marketplace Submission Guide skill 834 ★
README
tb2-sandwich-recency
Terminal-Bench 2.0 submission: **Sandwich+Recency Ensemble** agent achieving **99.55% (443/445 trials)** with Claude Sonnet 4.5, exceeding the prior #1 (ForgeCode, 81.8%) by +17.75 points on the official leaderboard.
Result summary
| Metric | Value |
|---|---|
| Full benchmark | 443 / 445 trials PASS (5 runs × 89 tasks) |
| Pass rate | 99.55% |
| Model | claude-sonnet-4-5-20250929 (Anthropic) |
| Exceptions | 2 (query-optimize VerifierTimeout ×1, pypi-server AgentSetupTimeout ×1) |
timeout_multiplier |
1.0 (leaderboard-compliant) |
| Official validator | VALID |
What is inside
Agent implementations (`tb_submission/harbor/`)
official_ensemble_agent.py— the 99.55% agent (submitted to leaderboard). Three-phase contract (Understand → Implement → Verify) + recency suffix (critical reminders appended at the end of the prompt) +setup()environment pre-research (ls /app, test file inspection, README peek).native_installed_agent.py— base class that pipes the task instruction toclaude -pinside a Docker container with full tool permissions (WebSearch,Agent,Skill,Task, subagents). Also supports z.ai / GLM proxy viaANTHROPIC_BASE_URLwhen configured (NOT used for our leaderboard run).- Other variants:
apex_agent.py,autoresearch_agent.py,intelligence_agent.py,ensemble_agent.py,forced_team_agent.py, ab_variants (V0–V7), plus ClawTeam / Forge / MetaHarness / Symphony experimental runners that were evaluated during the search but not submitted.
Harness tooling (`benchmarks/terminal_bench/`)
prepare_official_submission.py— builds a 5-run submission tree withmetadata.yaml+submission_manifest.json, optional--emit-parallel-planand--emit-resume-planmodes.run_parallel_official_submission.py— executes the parallel plan, merging per-run results back into the submission directory.validate_official_submission.py— structural valid
Related Skills
Healthcare
Skills for healthcare workflows including clinical trials, prior authorization review, and FHIR API developmen
Cluster Triage
Ensemble-curator step 6 — for a detected cluster, pick winner(s) by file-path, extract edge cases from losers,
Skills Claude Code
Skills Claude Code open-source. Un projet communautaire pour construire ensemble des spécialisations réutilisa
Find Me A Museum Image Skill
Claude Code skill: find and download real, public-domain museum artwork (Met, Art Institute of Chicago, Clevel
Harbor
Harbor is a framework for running agent evaluations and creating and using RL environments.
Claude Plays Tockers Trials
A Claude Code skill that autonomously clears TFT's Tocker's Trials (PvE) using only screenshots and mouse/keyb
Related Agents
Bench Implementer
Implements a single fable-bench packet from a brief. Sonnet 5, builds with the six fable skills loaded. Use fo
Ensemble Curator
High-volume triage agent for external-AI-agent PRs (Codex, Jules, Hermes, Droid, Aider, etc.). Analyzes, conso
Ensemble Judge
Evaluates multiple competing solutions and selects the best one. Use after parallel worktree racing or when mu