0.12.0 release notes → Rust core · Python native · ISO 8000/25012

Know your data before it bites.

Profile CSV, JSON, Parquet, DataFrames and Arrow in one call: types, nulls, distributions, patterns and a quality score. Then gate your pipeline on what it found. It runs locally, gives the same numbers however you load the data, and says “can’t tell” instead of guessing.

crates.io PyPI docs.rs MIT or Apache-2.0
python
import dataprof as dp

report = dp.profile("orders.csv")

report.rows, report.columns, report.quality_score
(200, 5, 92.98)

[(f.code, f.column) for f in report.findings()]
[('locale_numbers', 'price'),
 ('null_heavy', 'email'),
 ('sensitive_pattern', 'email')]

report.check(min_quality_score=90).verdict
'pass'  # or 'fail', or 'inconclusive'
The first ten minutes

Built for the questions you ask unfamiliar data

Find sparse columns, unstable types, duplicate keys, stale timestamps, and suspicious values before they turn into pipeline bugs.

Which columns are thin, empty, or structurally broken?

Null counts, completeness metrics, and schema shape in one pass.

Did this feed drift or spike somewhere suspicious?

Numeric summaries, outlier signals, and range checks.

Are these IDs really unique, or just pretending to be keys?

Distinct counts, uniqueness ratios, and duplicate warnings.

Are my timestamps plausible and fresh?

Future-date detection, stale-data signals, and timeliness scoring.

Did parsing silently go wrong?

Type inference, pattern matches, format violations, and source metadata.

What changed since yesterday?

Save a baseline report and diff it against today’s data, column by column.

Quickstart

Profile, then gate

A Python package for notebooks, scripts and CI, and a Rust crate for services and batch jobs. Same engine, same numbers.

uv pip install dataprof
import dataprof as dp

report = dp.profile("data.csv")          # files, dicts, bytes, DataFrames, Arrow
print(report.quality_summary())        # per-dimension ISO scores

result = report.check(min_quality_score=90, max_null_percentage={"*": 20})
print(result.verdict)                        # pass, fail or inconclusive

report.save("report.json")             # full report, reloadable
print(report.to_llm_context(max_tokens=500))  # for an agent, no raw values

In CI, with no code: python -m dataprof.check data.csv --min-quality 90 exits 0, 1 or 2 for pass, fail or inconclusive. Wheels for CPython 3.10 to 3.14 with no Python dependencies; pandas is optional, for DataFrame-typed exports. See the Python API guide.

What you can count on

Careful numbers, bounded memory

Same data, same numbers

CSV, Parquet, pandas, polars and Arrow of the same values produce the same profile on every engine. CI checks it on every push.

Bounded memory

Files larger than RAM stream through fixed-size accumulators, and exact counts hold up to a million distinct values.

Multi-format by default

CSV, JSON, JSONL, Parquet, live databases, DataFrames, and Arrow batches — one tool across all of them.

Verdicts that can say “can’t tell”

A sampled or partial scan never passes a claim about the whole file. Scores carry their confidence interval, and the gate decides on it.

Absence is not zero

A metric that was not computed is reported as missing, never as a plausible default. Findings list the rules they could not evaluate.

Agent-safe summaries

Token-bounded LLM context that carries counts and patterns, never your values. Profiling that plugs into agent workflows.

ISO 8000-8 · ISO/IEC 25012

A quality model you can point an auditor at

When quality analysis is requested, dataprof assesses up to seven dimensions informed by international data-quality standards. Its configurable aggregate score uses only the dimensions the data could actually support.

Completeness

Missing-cell percentage, share of fully-populated rows, columns past the null threshold.

Consistency

Data type consistency, format violations, encoding issues.

Uniqueness

Duplicate rows, key uniqueness, high-cardinality warnings.

Accuracy

Outlier ratio, range violations, negatives in positive-only columns.

Timeliness

Future dates, stale-data ratio, temporal ordering violations.

Validity

Conformance to confidently detected semantic patterns, with weak evidence left unassessed.

Precision

Consistency of observed decimal scale within floating-point columns.

Supported inputs

Meet your data where it lives

FormatEngineNotes
CSVIncremental, ColumnarAuto-detects , ; | \t delimiters
JSON / JSONLIncrementalArray-of-objects or one object per line
ParquetColumnarSchema and counts from metadata, no row scan needed
Database queryAsyncPostgreSQL, MySQL, SQLite via connection string
pandas / polars DataFrameColumnarPython API
Arrow RecordBatchColumnarZero-copy via PyCapsule, or the Rust API
dict / bytes / BytesIOColumnarPython API, no dependencies
Async byte streamIncrementalAny AsyncRead source (HTTP, WebSocket, …)
Measured, not promised

Every change is benchmarked in CI

Criterion runs on each push to master, and the full reports — throughput, scaling behavior, end-to-end pipeline timings — are published here, generated straight from the run’s artifacts.

Browse benchmark reports
Citing

Citing dataprof

dataprof has no associated publication or DOI, so the citation is for the software itself. GitHub’s Cite this repository button generates APA and BibTeX from CITATION.cff.

Reproducible benchmark material →

Contribute

Good first issues, with a mentor

A good first contribution here is small and self-contained: a focused test, a documentation fix, or one compiling example. Issues tagged mentor available come with someone to review and unblock you.

Pick an issue

Open tickets labelled good first issue, each with scope, acceptance criteria, and the commands to verify it.

Set up in one command

The contributor smoke path builds the extension and runs the focused tests:

python .github/scripts/contributor_smoke.py