GeoTheLeo

Music Man — Development skill for Claude Code

Development community

An agent that investigates music-streaming royalty fraud across millions of events and proposes withholding a payout — a human always approves before anything happens.

How to install Music Man

This entry records only its repository, not the path inside it, so there is no exact command to give. Open GeoTheLeo/music-man and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Music Man does

An agent that investigates music-streaming royalty fraud across millions of events and proposes withholding a payout — a human always approves before anything happens. PySpark + Delta Lake + MLflow lakehouse, Claude tool-calling agent.

Alternatives in Development

  • Arreflect — AgentRecall consolidation & reflection — periodic triage of recurring corrections; proposes rule changes, neve 370 ★
  • Cloud Run Debug — Diagnose a failing Cloud Run service — Antigravity (agy/Gemini) digests the error logs cheaply, Claude infers 284 ★
  • Music Genre Finder — 🎵 Intelligent music genre search with 5947 genres from RateYourMusic - Claude Code skill for quick lookup, sm 158 ★

README

Music Man — Music-Streaming Royalty-Fraud Lakehouse

An agent that investigates suspicious music-streaming patterns (bot-farm royalty fraud) and proposes withholding a payout — a human always approves before anything actually happens. Built on PySpark + Delta Lake + MLflow, running entirely locally, portable to a real Databricks workspace.

Companion project to [NorthStar's Agentic Intervention Copilot](../northstar-platform): same human-in-the-loop spine (agent proposes, never executes), different stack — this one proves out lakehouse-scale data engineering instead of a React/Next.js frontend.

Results (a real run, not a mockup)

Verified against the actual public catalog and a full pipeline run:

  • 89,740 real tracks, 31,428 real artists seeded from a public Kaggle Spotify dataset
  • 2.1M synthetic play events processed through the full bronze → silver → gold medallion pipeline on local PySpark + Delta Lake
  • 40 injected fraud cases across 4 distinct bot-farm signatures — the trained classifier caught 100% of them (recall 1.0) on a held-out test split, at 67% precision. The false positives are genuinely popular artists whose scale overlaps with the fraud signature — which is exactly the argument for why a human reviews every hold before it takes effect, not a flaw to hide.

**A real excerpt from a live agent run** — it didn't just call tools in sequence, it caught a problem in my own synthetic-data generator mid-investigation:

*"device_concentration_ratio is 0.0668 here and 0.0667–0.0669 for every other top candidate inspected (Cachureos, King 810, Feid). That near-identical value across unrelated artists looks like a pipeline artifact rather than an independent signal, so it should NOT be weighted as corroborating evidence."*

Unprompted, it also noticed the model's anomaly scores were saturated at 1.0 across 39 artists, refused to rank candidates by score alone, tie-broke on financial exposure instead, and flagged that the shared signa