andremiguel1

AI Native Playbook — Testing skill for Claude Code

Testing community

The harness I use to run coding agents on regulated production codebases: agent contracts, skills as runbooks, spec-driven delivery, memory, change management and evidence discipline.

How to install AI Native Playbook

This entry records only its repository, not the path inside it, so there is no exact command to give. Open andremiguel1/ai-native-playbook and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What AI Native Playbook does

The harness I use to run coding agents on regulated production codebases: agent contracts, skills as runbooks, spec-driven delivery, memory, change management and evidence discipline. Every rule carries the incident that created it.

Alternatives in Testing

README

[🇧🇷 Português (Brasil)](README.pt-BR.md) · 🇬🇧 **English**

AI-Native Engineering Playbook

**How I run coding agents on regulated production codebases without losing control.**

This is the operating system I use to ship software with AI coding agents (Claude Code, Codex, Copilot) on systems where mistakes cost money: an international payments platform (card-network messaging, real-time settlement, FX, AML reporting) and [Sinal](https://sinal.andrecristino.cloud), a multi-tenant AI-agent SaaS that I designed, build and operate. Every rule here exists because something went wrong without it.

**What is real and what is not.** The rules, the incidents and their dates are real. Client and partner names are absent on purpose, and so is anything operational: hostnames, account identifiers, volumes, totals. The worked example is fictional. Where a figure appears it is a ratio, or a count of my own artefacts. Never a client's business number.

Think of it as a harness. The contracts, runbooks, specs, memory and evidence discipline that let an agent do real work on a codebase without turning it into a liability. Prompts are the smallest part of it.

Two chapters go further than the harness, because there the agents are also the product. [Evaluating an LLM product](docs/09-evaluating-llm-products.md) covers the evaluation harness I built into a live agent SaaS: persona stress tests, LLM-as-judge against a rubric, runs keyed by the hash of the prompt so comparisons stay honest, hallucination scored apart from an honest gap. [Security hardening](docs/08-security-hardening.md) covers the seven-front audit method and the toolkit that *proves* a fix instead of assuming it. If you only read one of these ten documents, read [07. Evidence discipline](docs/07-evidence.md). It is the one that changes how you work tomorrow.

The model in one picture

            ┌──────────────────────────────────────────────────────────────┐
            │                       WORKING AGR