Persona Portability Benchmark
Description
One persona, one frozen memory, N models: how much of an agent's character survives a model swap? Harness + blind multi-lens judge panel + cross-vendor rank control + contamination probe. Results included: 7 models, spread 4.80 -> 2.20.
Installation
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open the source below and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
README
Persona Portability Benchmark
**One persona. One frozen memory. N models. How much of the character survives the model swap?**
Everyone versions their agent's prompt and memory. Almost nobody measures what happens to the agent's *character* when the model underneath changes. We could not find a public benchmark for "same persona on N models" (informal search, 2026-08-02 — prior art pointers welcome in issues), so we built one and ran it on our own working synthetic co-founder.
Result of our run (2026-08-03)
Same persona file, same frozen memory, byte-identical prompt envelope, 7 models, 5 character-probing tasks, blind 3-judge panel + cross-vendor rank control:
| model | score (1–5) |
|---|---|
| Claude Fable 5 | 4.80 |
| Claude Opus 5 | 4.60 |
| Claude Sonnet 5 | 4.00 |
| OpenAI Codex CLI (GPT) | 3.20 |
| Google Gemini CLI | 3.20 |
| Claude Haiku 4.5 | 2.80 |
| xAI Grok CLI | 2.20 |
Spearman rho between the Claude judge panel and an independent Gemini control judge: **0.839**, top-4 identical. Full tables, per-task breakdown, contamination report and limitations: [results/2026-08-03-run.md](results/2026-08-03-run.md).
Three takeaways:
- Character breaks before knowledge does. Weak models keep the facts (they hold the interrogation task) but lose calibration: they fold under emotional pressure, silently rewrite memory when the principal pushes, and replace a working scorecard with generic advice.
- It looks like a cliff, not a gradient (descriptive, one run: the two leaders sit 0.6+ above the field; the largest adjacent gap, 0.80, is right below third place). If your agent's model gets silently downgraded, memory stays — the character walks. We did not test the downgrade scenario directly; the ladder is the evidence.
- Stronger models fabricate better evidence. A persona that demands hard, evidence-based pushback converts, in strong models, into invented evidence (fake citations, invented grants, a non
Related Skills
Agency Agents
A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy inject
AI Awesome Llm Apps
100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
AI Firecrawl
🔥 The API to search, scrape, and interact with the web for AI
AI Artifacts Builder
Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech
AI Headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agen
AI CrewAI
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewA
AI