dataprof
Files, DataFrames, and Arrow streams. Quality reports that disclose what was actually assessed.
HELLO, WORLD. MAKE YOURSELF AT HOME.
Data engineer, open-source builder, and incurable tinkerer. I work where data pipelines, storage, and developer tools meet, usually with Rust, Python, or Go.
Have a look aroundTools I build, maintain, and learn from. The source is yours to explore.
Files, DataFrames, and Arrow streams. Quality reports that disclose what was actually assessed.
Collecting open data and turning messy sources into useful datasets.
Connect decisions, incidents, and the reasons things changed. Knowledge that stays in Git.
Runnable lakehouse pipelines, analytics marts, and quality gates. Also tested locally with DuckDB.
Supervised stream ingestion, durable events, and reproducible edge benchmarks.
From the workbench, summer 2026: Lares · Fantabuddy · OCCAS · Iceberg commit experiment
Research & reproducible benchmarks: Nephtys / UIC 2026 · dataprof / ScalCom 2026
I contribute fixes and improvements to the tools I use, including Apache Arrow, DataFusion, Iceberg Rust, Polars, and Tokio.
Follow the contribution trail →Experiments, things that broke, and what I learned along the way.
Data infrastructure, open source, or a good technical rabbit hole. Drop me a line.