Onboarding Clean — Data skill for Claude Code
Write and run data cleaning code to produce partitioned parquet output.
How to install Onboarding Clean
Installs to ~/.claude/skills/basedosdados-pipelines-onboarding-clean/SKILL.md
mkdir -p ~/.claude/skills/basedosdados-pipelines-onboarding-clean && curl -fsSL https://raw.githubusercontent.com/basedosdados/pipelines/HEAD/.claude/skills/onboarding-clean.md -o ~/.claude/skills/basedosdados-pipelines-onboarding-clean/SKILL.md Restart Claude Code, or start a new session, for it to be picked up.
What Onboarding Clean does
description: Write and run data cleaning code to produce partitioned parquet output
argument-hint: [output_path]
Spawn the `cleaner` agent with: $ARGUMENTS
The agent will read architecture tables from Drive, inspect the raw files, write Python cleaning code, validate a subset, and scale to full data after user confirmation.
Alternatives in Data
- /recon — Run the full recon pipeline on a target and produce a prioritized attack surface 927 ★
- Sprite Gen — Generate clean 2D game sprites & animation atlases — component-row pipeline: state rows, alpha cleanup, frame 763 ★
- Reddit MCP Buddy — Clean, LLM-optimized Reddit MCP server 624 ★
Full documentation available on GitHub
View Source RepositoryRelated Skills
00 Process Video
Run the full video editing pipeline on a raw talking-head video to produce a polished final output, with optio
Onboarding Dbt
Write DBT .sql and schema.yml files for a Data Basis dataset
07 Clean Artifacts
Remove intermediate video files from the pipeline, keeping only the original and final output
Fix Tickets
Autonomously fix one or more speckit-companion GitHub issues with the SpecKit Companion pipeline — clean branc
Deai
deai — de-AI writing for Chinese text. One self-contained Agent Skill: strip AI flavor from web novels, fictio
Multi AI Session Data Extractor
Capture and preserve your own AI conversations locally before they vanish — 10 platforms, unified parquet sche
Related Agents
Foley
Backlog triage analyst — stage 1 of the backlog pipeline. Invoke with pre-fetched task data to assess dispatch
Shine Data Engineer
Local data analysis via DuckDB/SQLite MCP — SQL on CSV/Parquet/JSON/Excel without cloud services. Produces cha
Data Reviewer
Reviews datasets, schemas, historian tag lists, MES/ERP tables, PI/AF trees, CSV/Parquet exports, Kafka topics