AI Native Playbook — Testing skill for Claude Code
The harness I use to run coding agents on regulated production codebases: agent contracts, skills as runbooks, spec-driven delivery, memory, change management and evidence discipline.
How to install AI Native Playbook
This entry records only its repository, not the path inside it, so there is no
exact command to give. Open andremiguel1/ai-native-playbook and copy the folder into
~/.claude/skills/, or the file into ~/.claude/agents/.
What AI Native Playbook does
The harness I use to run coding agents on regulated production codebases: agent contracts, skills as runbooks, spec-driven delivery, memory, change management and evidence discipline. Every rule carries the incident that created it.
Alternatives in Testing
- Deliver — Delivery phase - Review, validate, and test with multi-AI quality assurance 2.8k ★
- AI Contribution Skills (Engram) — branch-pr: clean branch and PR workflow 2.5k ★
- Test Driven Development.SKILL — Enforces TDD discipline with RED-GREEN-REFACTOR cycle 262 ★
README
[🇧🇷 Português (Brasil)](README.pt-BR.md) · 🇬🇧 **English**
AI-Native Engineering Playbook
**How I run coding agents on regulated production codebases without losing control.**
This is the operating system I use to ship software with AI coding agents (Claude Code, Codex, Copilot) on systems where mistakes cost money: an international payments platform (card-network messaging, real-time settlement, FX, AML reporting) and [Sinal](https://sinal.andrecristino.cloud), a multi-tenant AI-agent SaaS that I designed, build and operate. Every rule here exists because something went wrong without it.
**What is real and what is not.** The rules, the incidents and their dates are real. Client and partner names are absent on purpose, and so is anything operational: hostnames, account identifiers, volumes, totals. The worked example is fictional. Where a figure appears it is a ratio, or a count of my own artefacts. Never a client's business number.
Think of it as a harness. The contracts, runbooks, specs, memory and evidence discipline that let an agent do real work on a codebase without turning it into a liability. Prompts are the smallest part of it.
Two chapters go further than the harness, because there the agents are also the product. [Evaluating an LLM product](docs/09-evaluating-llm-products.md) covers the evaluation harness I built into a live agent SaaS: persona stress tests, LLM-as-judge against a rubric, runs keyed by the hash of the prompt so comparisons stay honest, hallucination scored apart from an honest gap. [Security hardening](docs/08-security-hardening.md) covers the seven-front audit method and the toolkit that *proves* a fix instead of assuming it. If you only read one of these ten documents, read [07. Evidence discipline](docs/07-evidence.md). It is the one that changes how you work tomorrow.
The model in one picture
┌──────────────────────────────────────────────────────────────┐
│ WORKING AGR
Related Skills
Agent Runway
Operational discipline layer for AI-assisted software delivery: spec/ticket workflows, quality gates, and pers
Harness Agentic AI Agent Best Practices And Use Case
Production-ready AI agent for UI/web testing on Amazon Bedrock AgentCore Harness — Memory, Skills, observabili
Alih Spec
⚡ Enterprise Spec-Driven Development (SDD) framework & native AI skill for converting codebases across stacks.
AI Native Sdlc Skills
Unofficial Claude Code skills adapted from Anthropic's The AI-Native SDLC playbook — the committed-artifact ch
Claude Code Multi Agent Harness
Cursor Agent Skill: spec-delegate-review harness, sourcing discipline, academic slice
Check Contracts
Verify API contracts in /contracts/ match actual implementations. Detects drift between spec and code before i
Related Agents
Planner Architect
Analyzes codebases, designs architecture, decomposes complex requests into parallel tasks, generates interface
Duplicate Rule Designer
Design a Matching Rule + Duplicate Rule pair. SfSkills data workflow agent, /design-duplicate-rule: reads its
Code Security Auditor
Comprehensive security analysis and vulnerability detection for codebases. Specializes in threat modeling, sec