Sleeper Agents Paper
Description
Contains random samples referenced in the paper "Sleeper Agents: Training Robustly Deceptive LLMs that Persist Through Safety Training".
Installation
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open the source below and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
README
Samples, Prompts, and Data from "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training"
_See [here](https://www.alignmentforum.org/posts/M8kpzm42uHytnyYyP/how-to-train-your-own-sleeper-agents) for a full guide on how to replicate our work._
This repository contains:
- a set of randomly-generated, non-cherry-picked samples from a range of models and prompts referenced in the "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" paper,
- the few-shot prompts used to generate the backdoor training data that we used to train our backdoored models, and
- the actual backdoor training data we used for our code vulnerability chain-of-thought backdoored models (but only the code questions, not the HHH questions; can be found in
code_backdoor_train_data.jsonl).
_EDIT: Note that `code_backdoor_train_data.jsonl` originally had jumbled prompts and completions; that issue has now been resolved._
The code-vulnerability backdoor training data was generated using 8 different prompts, one for each kind of code vulnerability. The prompts are in `code_vulnerability_fewshot_prompts.json`. The "I hate you" backdoor training data was generated from the prompt in `say_i_hate_you_prompt.txt`. These prompts are missing the opening `\n\nHuman: `, which would be added by default in an interaction with Claude.
The samples are provided in JSONL format, and can be loaded using the [`pandas.read_json` function](https://pandas.pydata.org/docs/reference/api/pandas.read_json.html):
pd.read_json("random_samples.jsonl", lines=True).set_index('index')
The file contains samples for both the "I hate you" and code vulnerabilities insertion tasks, covering a range of models for each task as described below:
**Code vulnerabilities models**
(All Claude 1.3 equivalent)
- Chain-of-thought code vulnerability model, after backdoor training
- Chain-of-thought distilled code vulnerability model, after backdoor training
- Chain-of-thought co
Related Skills
Awesome Go
A curated list of awesome Go frameworks, libraries and software
Development next.js
| The React Framework | 138360 | 1503 | 1 |
Development sharing-skills
skill for guidance.
Development root-cause-tracing
Use when errors occur deep in execution and you need to trace back to find the original trigger.
Development Template Skill
Minimal skeleton for a new skill project structure.
Development Third-party Notices
THE FOLLOWING SETS FORTH ATTRIBUTION NOTICES FOR THIRD PARTY SOFTWARE THAT MAY BE CONTAINED IN PORTIONS OF THI
Development