Deep Rl — Development skill for Claude Code
Deep reinforcement learning - DQN/Rainbow/R2D2/Agent57/BBF, PPO/TRPO/GRPO, SAC/TD3/REDQ/DroQ/CrossQ, DreamerV3/TD-MPC2/MuZero, CQL/IQL/TD3+BC/AWAC/Decision Transformer, MAPPO/IPPO multi-agent, Go-Expl.
How to install Deep Rl
Installs to ~/.claude/commands/tachyon-beep-skillpacks-deep-rl.md
mkdir -p ~/.claude/commands && curl -fsSL https://raw.githubusercontent.com/tachyon-beep/skillpacks/HEAD/.claude/commands/deep-rl.md -o ~/.claude/commands/tachyon-beep-skillpacks-deep-rl.md Restart Claude Code, or start a new session, for it to be picked up.
What Deep Rl does
description: Deep reinforcement learning - DQN/Rainbow/R2D2/Agent57/BBF, PPO/TRPO/GRPO, SAC/TD3/REDQ/DroQ/CrossQ, DreamerV3/TD-MPC2/MuZero, CQL/IQL/TD3+BC/AWAC/Decision Transformer, MAPPO/IPPO multi-agent, Go-Explore/NGU/BYOL-Explore, reward shaping, counterfactual reasoning (HER/OPE), debugging, evaluation. Routes to 13 specialist sheets, 3 commands, 2 SME agents.
Deep RL Routing
**Problem type determines algorithm family. RL is not one algorithm - action space (discrete vs continuo
Alternatives in Development
- PaperSpine — PaperSpine is a motivation-driven skill for learning from strong academic papers, building a paper’s central a 5k ★
- Claudeception — A Claude Code skill for autonomous skill extraction and continuous learning 2.2k ★
- Cq — An open standard for shared agent learning 1.3k ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
DeepTerrainRL
terrain-adaptive locomotion skills using deep reinforcement learning
PARL
PARL (Parallel-Agent Reinforcement Learning) is a training paradigm that teaches models to decompose complex t
InstinctMJ
Provide a native mjlab environment for Project-Instinct to support reinforcement learning in humanoid whole-bo
Claudehut Learning Report
Show the ClaudeHut learning scoreboard — measured memory health (store size, reinforcement, effectiveness/recu
Compound Learning
Compound Learning & Deep Planning
Twin Sparrow Deep Structure Learning
Enter deep-structure learning mode for books, PDFs, passages, theories, symbols, dreams, religions, scientific
Related Agents
Deep Rl Algorithms Curator
Curator for the deep RL algorithms lane of Reinforcement Learning Brain. Use when maintaining source coverage,
Polymathic Gates
Reasons through Bill Gates's cognitive architecture — platform thinking, systematic deep-reading, multi-source
Robo AI
Robot learning and perception — vision pipelines, 3D detection/segmentation, pose estimation, VLA and robot fo