anthony-maio

Triton GPU Kernel Optimization — Development skill for Claude Code

Development community

Write high-performance tiled Triton GPU kernels with autotune, grouped tile ordering, stride-based addressing, and proper benchmarking.

How to install Triton GPU Kernel Optimization

Installs to ~/.claude/skills/anthony-maio-triton-skills-triton-gpu-kernel-optimization/SKILL.md

Terminal
mkdir -p ~/.claude/skills/anthony-maio-triton-skills-triton-gpu-kernel-optimization && curl -fsSL https://raw.githubusercontent.com/anthony-maio/triton-skills/HEAD/skills/triton-gpu-kernel-optimization.md -o ~/.claude/skills/anthony-maio-triton-skills-triton-gpu-kernel-optimization/SKILL.md

Restart Claude Code, or start a new session, for it to be picked up.

What Triton GPU Kernel Optimization does


name: triton-gpu-kernel-optimization description: Write high-performance tiled Triton GPU kernels with autotune, grouped tile ordering, stride-based addressing, and proper benchmarking.

Tiled GEMM & General Kernel Optimization in Triton

**Targets:** Triton >= 2.1, SM70+/CDNA2+

Overview

High-performance Triton kernels follow a consistent structure: block-tiled work distribution, stride-based pointer arithmetic, FP32 accumulation, boundary masking, and autotune sweeps. This file

Alternatives in Development

  • Claudes C Compiler — Claude Opus 4.6 wrote a dependency-free C compiler in Rust, with backends targeting x86 (64- and 32-bit), ARM 2.6k ★
  • Fable OS — An agentic operating system where the kernel is controlled directly by Claude 324 ★
  • Obsidian Health — Run a vault health check — grouped by severity, detects contradictions, concept gaps, stale claims, and struct 277 ★

Full documentation available on GitHub

View Source Repository