Which columns are thin, empty, or structurally broken?
Null counts, completeness metrics, and schema shape in one pass.
Profile CSV, JSON, Parquet, DataFrames and Arrow in one call: types, nulls, distributions, patterns and a quality score. Then gate your pipeline on what it found. It runs locally, gives the same numbers however you load the data, and says “can’t tell” instead of guessing.
import dataprof as dp
report = dp.profile("orders.csv")
report.rows, report.columns, report.quality_score
(200, 5, 92.98)
[(f.code, f.column) for f in report.findings()]
[('locale_numbers', 'price'),
('null_heavy', 'email'),
('sensitive_pattern', 'email')]
report.check(min_quality_score=90).verdict
'pass' # or 'fail', or 'inconclusive'
Find sparse columns, unstable types, duplicate keys, stale timestamps, and suspicious values before they turn into pipeline bugs.
Null counts, completeness metrics, and schema shape in one pass.
Numeric summaries, outlier signals, and range checks.
Distinct counts, uniqueness ratios, and duplicate warnings.
Future-date detection, stale-data signals, and timeliness scoring.
Type inference, pattern matches, format violations, and source metadata.
Save a baseline report and diff it against today’s data, column by column.
A Python package for notebooks, scripts and CI, and a Rust crate for services and batch jobs. Same engine, same numbers.
uv pip install dataprof
import dataprof as dp
report = dp.profile("data.csv") # files, dicts, bytes, DataFrames, Arrow
print(report.quality_summary()) # per-dimension ISO scores
result = report.check(min_quality_score=90, max_null_percentage={"*": 20})
print(result.verdict) # pass, fail or inconclusive
report.save("report.json") # full report, reloadable
print(report.to_llm_context(max_tokens=500)) # for an agent, no raw values
In CI, with no code: python -m dataprof.check data.csv --min-quality 90
exits 0, 1 or 2 for pass, fail or inconclusive.
Wheels for CPython 3.10 to 3.14 with no Python dependencies; pandas is optional, for DataFrame-typed exports.
See the Python API guide.
cargo add dataprof
use dataprof::Profiler;
let report = Profiler::new().analyze_file("data.csv")?;
println!("Rows: {}", report.execution.rows_processed);
println!("Quality: {:.1}%", report.quality_score().unwrap_or(0.0));
for col in &report.column_profiles {
println!("{} {:?} nulls={}", col.name, col.data_type, col.null_count);
}
MSRV 1.96. Feature flags cover Arrow, Parquet, async streaming, and
PostgreSQL / MySQL / SQLite connectors — or go lean with
default-features = false.
See docs.rs.
CSV, Parquet, pandas, polars and Arrow of the same values produce the same profile on every engine. CI checks it on every push.
Files larger than RAM stream through fixed-size accumulators, and exact counts hold up to a million distinct values.
CSV, JSON, JSONL, Parquet, live databases, DataFrames, and Arrow batches — one tool across all of them.
A sampled or partial scan never passes a claim about the whole file. Scores carry their confidence interval, and the gate decides on it.
A metric that was not computed is reported as missing, never as a plausible default. Findings list the rules they could not evaluate.
Token-bounded LLM context that carries counts and patterns, never your values. Profiling that plugs into agent workflows.
When quality analysis is requested, dataprof assesses up to seven dimensions informed by international data-quality standards. Its configurable aggregate score uses only the dimensions the data could actually support.
Missing-cell percentage, share of fully-populated rows, columns past the null threshold.
Data type consistency, format violations, encoding issues.
Duplicate rows, key uniqueness, high-cardinality warnings.
Outlier ratio, range violations, negatives in positive-only columns.
Future dates, stale-data ratio, temporal ordering violations.
Conformance to confidently detected semantic patterns, with weak evidence left unassessed.
Consistency of observed decimal scale within floating-point columns.
| Format | Engine | Notes |
|---|---|---|
| CSV | Incremental, Columnar | Auto-detects , ; | \t delimiters |
| JSON / JSONL | Incremental | Array-of-objects or one object per line |
| Parquet | Columnar | Schema and counts from metadata, no row scan needed |
| Database query | Async | PostgreSQL, MySQL, SQLite via connection string |
| pandas / polars DataFrame | Columnar | Python API |
| Arrow RecordBatch | Columnar | Zero-copy via PyCapsule, or the Rust API |
| dict / bytes / BytesIO | Columnar | Python API, no dependencies |
| Async byte stream | Incremental | Any AsyncRead source (HTTP, WebSocket, …) |
Criterion runs on each push to master, and the full reports —
throughput, scaling behavior, end-to-end pipeline timings — are published here,
generated straight from the run’s artifacts.
dataprof has no associated publication or DOI, so the citation is for the software itself. GitHub’s Cite this repository button generates APA and BibTeX from CITATION.cff.
A good first contribution here is small and self-contained: a focused test, a documentation fix, or one compiling example. Issues tagged mentor available come with someone to review and unblock you.
Open tickets labelled good first issue, each with scope, acceptance criteria, and the commands to verify it.
The contributor smoke path builds the extension and runs the focused tests:
python .github/scripts/contributor_smoke.py
Usage questions and ideas belong in GitHub Discussions. The contributing guide covers conventions and review.