Triton Quantized Block Scaled Gemm
Description
--- name: triton-quantized-block-scaled-gemm description: Teach an AI agent how to implement block-scaled (microscaling) quantized matmul kernels in Triton. --- # Quantized & Block-Scaled Matmul Kernels in Triton > **Targets:** Triton >= 3.0; `tl.dot_scaled` requires SM100+/CDNA4; dequantize fallback works on SM70+/CDNA2+ Overview This guide explains how to implement low-precision block-scaled matrix multiplication in Triton for mxfp4/mxfp8/nvfp4 formats. It covers scale tensor layouts (OCP m
Installation
Installs to ~/.claude/skills/anthony-maio-triton-skills-triton-quantized-block-scaled-gemm/SKILL.md
mkdir -p ~/.claude/skills/anthony-maio-triton-skills-triton-quantized-block-scaled-gemm && curl -fsSL https://raw.githubusercontent.com/anthony-maio/triton-skills/HEAD/skills/triton-quantized-block-scaled-gemm.md -o ~/.claude/skills/anthony-maio-triton-skills-triton-quantized-block-scaled-gemm/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
Full documentation available on GitHub
View Source RepositoryRelated Skills
Agency Agents
A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy inject
AI Firecrawl
🔥 The API to search, scrape, and interact with the web for AI
AI Artifacts Builder
Suite of tools for creating elaborate, multi-component claude.ai HTML artifacts using modern frontend web tech
AI Headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agen
AI CrewAI
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewA
AI TrendRadar
⭐AI-driven public opinion & trend monitor with multi-platform aggregation, RSS, and smart alerts.🎯 告别信息过载,你的
AI