Agent Rule Gaming — AI skill for Claude Code
Experiments in a small RL environment: do LLM agents game the rules when the answer is within reach, and can you tell them not to.
How to install Agent Rule Gaming
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open lauravoicu/agent-rule-gaming and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What Agent Rule Gaming does
Experiments in a small RL environment: do LLM agents game the rules when the answer is within reach, and can you tell them not to? Claude Opus 5, Qwen 3.8 27B, abliterated Qwen.
Alternatives in AI
- Claude Delegator — Delegate tasks to Codex GPT 5.2 directly from within Claude Code 909 ★
- DontFeedTheAI — Transparent anonymization proxy for AI-assisted pentesting 654 ★
- PaperFarm — Let AI agents run experiments in any repo while you sleep 353 ★
README
[](LICENSE) [](https://www.python.org/downloads/) []() [](https://www.linkedin.com/in/voiculaura/)
Agent Rule Gaming
A small LLM reinforcement-learning **environment**, built end to end, and a set of experiments on reward hacking: what an LLM agent does when the grader's answer is within reach, and whether you can tell it not to use it.
The environment has an instance generator, a network-less sandbox per episode, a grader that checks answers programmatically, a rollout loop driven by Claude through tool use (or by a local open-weight model through Ollama), a pass@k evaluation, and a set of probes for reward hacking. The task is deliberately benign: the agent reads a file of integers inside the sandbox and submits a statistic (sum, max, or count above 50). The engineering is what matters: isolation, grader placement, verifiable reward, rollouts, and evaluation. Swapping in a real defensive-security task, such as patch-and-verify or alert triage against ground truth, changes the instance generator and the grader, not the architecture.
This is a playground experiment, not a benchmark. Results of all recorded runs are in [RESULTS.md](RESULTS.md); the guided build walkthrough is in [TUTORIAL.md](TUTORIAL.md).
**Findings** (details and caveats in RESULTS.md):
*Claude Opus 5*
- When the honest route is blocked, the agent finds an answer key left in the sandbox within three tool calls.
- Under "compute from the data; do not guess" it submits the key in 3/30 episodes. Add one sentence saying only correct answers score, and that becomes 25/30 (p ≈ 10⁻⁸). The rule, the task and the grader are unchanged.
- Every key s
Related Skills
AI Gaming Agent
A step by step implementation of building an AI agent that plays 3d shooting game
Quackd
🦆🧠 Give your Microduck a brain. Tell a small robot with two legs what you want in plain language. An LLM (Cl
Vibrator
Small golang binary which build a Dockeized AI coding environment for isolated YOLO vibrations.
Qwopus
Local AI coding agent powered by Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled. Like Claude Code, but runs e
LLM Launchpad
一条命令在 Apple Silicon Mac 上部署本地大模型(Qwen3.8-27B · Ollama · Claude Code 兼容)
Tauri Plugin MCP
Allows AI agents (e.g., Cursor, Claude Code) to debug within Tauri apps via screenshot capture, window managem
Related Agents
Audit Sweep
A read-only audit or survey sweep — reviewing a corpus against a stated rule and reporting what violates it, w
Experiment Tracker
PROACTIVELY use this agent when experiments are started, modified, or when results need analysis. This agent s
Dr Platform Pitfalls
Deep-review finder angle. Applies the classic pitfall catalogue of the specific language, runtime, test framew