Persona Portability Benchmark banner
tonydzi tonydzi

Persona Portability Benchmark

AI community

Description

One persona, one frozen memory, N models: how much of an agent's character survives a model swap? Harness + blind multi-lens judge panel + cross-vendor rank control + contamination probe. Results included: 7 models, spread 4.80 -> 2.20.

Installation

This entry records only its repository, not the path inside it, so there is no exact command to give. Open the source below and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

README

Persona Portability Benchmark

**One persona. One frozen memory. N models. How much of the character survives the model swap?**

Everyone versions their agent's prompt and memory. Almost nobody measures what happens to the agent's *character* when the model underneath changes. We could not find a public benchmark for "same persona on N models" (informal search, 2026-08-02 — prior art pointers welcome in issues), so we built one and ran it on our own working synthetic co-founder.

Result of our run (2026-08-03)

Same persona file, same frozen memory, byte-identical prompt envelope, 7 models, 5 character-probing tasks, blind 3-judge panel + cross-vendor rank control:

model score (1–5)
Claude Fable 5 4.80
Claude Opus 5 4.60
Claude Sonnet 5 4.00
OpenAI Codex CLI (GPT) 3.20
Google Gemini CLI 3.20
Claude Haiku 4.5 2.80
xAI Grok CLI 2.20

Spearman rho between the Claude judge panel and an independent Gemini control judge: **0.839**, top-4 identical. Full tables, per-task breakdown, contamination report and limitations: [results/2026-08-03-run.md](results/2026-08-03-run.md).

Three takeaways:

  1. Character breaks before knowledge does. Weak models keep the facts (they hold the interrogation task) but lose calibration: they fold under emotional pressure, silently rewrite memory when the principal pushes, and replace a working scorecard with generic advice.
  2. It looks like a cliff, not a gradient (descriptive, one run: the two leaders sit 0.6+ above the field; the largest adjacent gap, 0.80, is right below third place). If your agent's model gets silently downgraded, memory stays — the character walks. We did not test the downgrade scenario directly; the ladder is the evidence.
  3. Stronger models fabricate better evidence. A persona that demands hard, evidence-based pushback converts, in strong models, into invented evidence (fake citations, invented grants, a non