katelynltensley

Database Cleanup Tool — Data skill for Claude Code

Data community

Dedupes and enriches messy real estate agent contact databases: merges duplicates, fills in missing phone numbers and emails, and outputs a clean CSV plus a data-quality report.

How to install Database Cleanup Tool

This entry records only its repository, not the path inside it, so there is no exact command to give. Open katelynltensley/database-cleanup-tool and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

What Database Cleanup Tool does

Dedupes and enriches messy real estate agent contact databases: merges duplicates, fills in missing phone numbers and emails, and outputs a clean CSV plus a data-quality report. Python CLI, Claude-assisted matching.

Alternatives in Data

  • Context Mode — Benchmark Results — Benchmarked against real outputs from popular Claude Code MCP servers, Skills, and dev tools 5.6k ★
  • Sprite Gen — Generate clean 2D game sprites & animation atlases — component-row pipeline: state rows, alpha cleanup, frame 763 ★
  • Chart Design — Design a chart or data visualization — selects the right chart type, applies accessible color palettes, adds a 309 ★

README

Database Cleanup Tool

A command-line tool that takes a messy real estate agent contact export, merges the duplicates, fills in missing phone numbers and emails from whatever the database already has scattered across those duplicates, and optionally runs whatever's still missing through a real paid lookup provider. It outputs a clean CSV and a data-quality report explaining exactly what changed and why.

Built to run unattended — one command, no prompts, no manual review step required to finish a run — so it can be scheduled instead of run by hand.

Why this exists

Real estate agent databases accumulate years of duplicate, partial, and inconsistently formatted contacts: the same person entered twice with different phone formatting, a nickname instead of a given name, an old address that's been updated on one entry but not the other. Rule-based dedupe tools tend to miss the near-matches that matter (a nickname plus a typo'd address) and flag ones that don't (two different people who happen to share a last name). This tool splits the problem the way it should be split: deterministic code handles everything that's actually deterministic — phone/email formatting, exact matches, clearly-not-a-match pairs — and Claude is only called in for the genuinely ambiguous middle, capped so a 10,000-row database can't generate a surprise bill.

What it actually does, in order

  1. Parse — flexible column-name mapping, so it isn't locked to one CRM's export format (dbcleanup/parser.py).
  2. Normalize — phone numbers, emails, names, and addresses get put into consistent formats so later comparisons are apples-to-apples (dbcleanup/normalize.py).
  3. Match — exact phone/email matches and clear fuzzy matches (name + address) are resolved by rules alone. Only pairs that land in the ambiguous middle get sent to Claude for a same-person judgment call, up to a configured budget; beyond that budget (or with Claude turned off) a conservative fallback