dataprof / performance lab

See how dataprof performs.

From a CSV scan to a complete profile. Explore measured workloads, inspect their variation, and run the same experiments on your own data infrastructure.

01 / Python tool comparison

One input. Explicit workloads.

File-to-summary timings for 1,000 rows of numeric, text and null data. Every sample checks row counts, column order and nulls.

3 samples per condition · 1 process blocks
Wall time · median with interquartile range

Warm operations

Repeated CSV reads and fresh summaries after warmup. Adapter import/setup excluded.

dataprof 1.66 ms
pandas 3.37 ms
polars 0.97 ms
ydata-profiling 216.92 ms
● Median   ▰ Middle 50% of samples (IQR) Log scale from 0.61 ms to 353.87 ms · Lower is faster

Fresh-process operations

A fresh process per sample. Startup, imports, operation and exit included.

dataprof 78.55 ms
pandas 359.79 ms
polars 154.89 ms
ydata-profiling 2.61 s
● Median   ▰ Middle 50% of samples (IQR) Log scale from 57.41 ms to 3.60 s · Lower is faster
Read the workload before the ratio. Each tool computes a different summary; metric equivalence is not established. Fresh process does not imply cold storage. These measurements describe this run, not a universal ranking. IQR shows variation, not a confidence interval.
How to read these numbers

Diagnostic evidence — no established baseline. Publication mode increases the experiment budget; it does not certify precision. Warm samples share a process within each block. Separate blocks use fresh processes, but share host and OS caches. Within-run intervals do not demonstrate across-run stability; repeat-run IQR overlap is a diagnostic, not a significance test.

Optional resource collection enabled. Energy and peak RSS are per worker, including imports and warmups; warm values cover the entire block, not a single operation. Host energy zones include background load and collector overhead and are never summed. Unavailable counters remain unavailable. See the resource median/IQR tables and raw readings and paired idle baselines. Shared-runner results are collection smoke evidence, not energy-efficiency claims.

Ordered samples and timing boundaries

Seconds in execution order; no first sample is discarded. Process time includes interpreter startup, harness imports, adapter import/setup, operations, validation, garbage collection, IPC and exit. Import/setup is timed directly around the adapter; deferred imports remain inside operations. First operation includes initialization and, for warm workers, is the first warmup. Warmup samples are retained in the raw JSON.

First adapter import after environment preparation (preflight): dataprof: 0.009436; pandas: 0.198340; polars: 0.057147; ydata-profiling: 1.690976. Earlier installs or workflow preflights may have warmed library pages; this is not disk-cold.

Ordered worker observations (seconds)
OrderBlockToolModeInvocation ProcessImport/setupFirst operationMeasured operations
11polarscoldfirst_fixture_operation_after_preflight0.1614290.0575060.0030900.003090
21pandascoldfirst_fixture_operation_after_preflight0.3875400.2168970.0054930.005493
31dataprofcoldfirst_fixture_operation_after_preflight0.0796030.0098140.0049570.004957
41ydata-profilingcoldfirst_fixture_operation_after_preflight2.8016961.6796540.3730610.373061
51dataprofcoldsubsequent_fresh_process0.0785460.0094160.0049080.004908
61pandascoldsubsequent_fresh_process0.3596840.2071510.0051740.005174
71ydata-profilingcoldsubsequent_fresh_process2.6120881.6497410.2069730.206973
81polarscoldsubsequent_fresh_process0.1548860.0593410.0031090.003109
91polarscoldsubsequent_fresh_process0.1531010.0573960.0030600.003060
101dataprofcoldsubsequent_fresh_process0.0741950.0091020.0047860.004786
111pandascoldsubsequent_fresh_process0.3597920.2088400.0049430.004943
121ydata-profilingcoldsubsequent_fresh_process2.5492701.5827320.2116120.211612
131dataprofwarmsubsequent_fresh_process0.0931220.0098680.0048950.001664, 0.001682, 0.001583
141pandaswarmsubsequent_fresh_process0.3980730.2074120.0051650.003541, 0.003375, 0.003333
151ydata-profilingwarmsubsequent_fresh_process3.6445221.5627680.3179270.211409, 0.216925, 0.239258
161polarswarmsubsequent_fresh_process0.1590080.0569750.0029100.001042, 0.000968, 0.000916

Minimal-worker controls, before each round (seconds): block 1, round 1: 0.017430; block 1, round 2: 0.017411; block 1, round 3: 0.016300.

Control scope: interpreter, minimal JSON worker, IPC and exit, without harness or tool imports. It is context for combined startup costs, not a subtractable import estimate.

Exact workloads and timing table
Seconds: median [IQR]
ToolCold median [IQR]Warm median [IQR] Workload
dataprof0.078546 [0.002704]0.001664 [0.000049]profile(engine="auto", metrics=["schema", "statistics"])
pandas0.359792 [0.013928]0.003375 [0.000104]read_csv + describe(include="all") + null counts
polars0.154886 [0.004164]0.000968 [0.000063]read_csv + describe + null counts
ydata-profiling2.612088 [0.126213]0.216925 [0.013924]pandas.read_csv + ProfileReport(minimal=True).description_set
Machine, versions and fixture identity
Measured at
2026-10-08T07:57:22.883260+00:00
Operating system
Linux-6.17.0-1022-azure-x86_64-with-glibc2.39
Processor
AMD EPYC 9V45 96-Core Processor
Samples / warmups per block
3 / 1
Process blocks
1
Experiment mode
routine
Declared host controls
GitHub-hosted runner; resource collection smoke test only
Library cache
not evicted; environment preparation and preflight may warm library pages and caches
Import preflight
imports only, in disposable workers before measurements; no fixture operations; may warm OS library pages and on-disk library caches, which are not evicted; fresh-process samples are subsequent to preflight, not first host invocations
Thread request
1
Cold file cache
warm (residency / eviction unverified)
Fixture SHA-256
1d766b0c21f3a89d78c31e76c9196471a2b5c7a87e6c39bdcff87e44b8e42272
Checkout
85c0850eb934a9f1277231a38129626297e0991c
Working tree
clean
Installed tools
dataprof 0.12.0, pandas 2.3.3, polars 1.44.2, ydata-profiling 4.18.4
Python / Arrow handoff

Python / Arrow boundaries

What it costs to hand pandas, polars and Arrow data to dataprof, stage by stage.

How to read these numbers

Diagnostic evidence, no established baseline. Exact serialized column metrics, absence, integers, nulls and order are checked before comparison. Consumer import and profiling share one timer; lazy stream construction happens inside it. End-to-end includes both independent exports. Fresh means a new process, not cold storage.

Peak RSS includes native buffers, imports and warmups across the worker lifetime. Arrow pool observations are partial allocator evidence, not profiler-only memory.

Stage times: median [IQR], seconds
Producer / rows / chunk / offset ConditionPrepareImport + profileDictJSON End-to-end
arrow_array/1000/250/offset-0Skipped: C Array exports one batch; chunked input uses Table or C Stream
arrow_array/1000/250/offset-3Skipped: C Array exports one batch; chunked input uses Table or C Stream
arrow_array/1000/1000/offset-0fresh0.135536 [0.003819]0.005084 [0.000562]0.000118 [0.000009]0.000053 [0.000014]0.140791 [0.003280]
arrow_array/1000/1000/offset-0warm0.000492 [0.000003]0.001317 [0.000018]0.000085 [0.000001]0.000040 [0.000000]0.001934 [0.000022]
arrow_array/1000/1000/offset-3fresh0.136664 [0.000184]0.004533 [0.000049]0.000123 [0.000006]0.000041 [0.000001]0.141361 [0.000240]
arrow_array/1000/1000/offset-3warm0.000485 [0.000016]0.001346 [0.000024]0.000081 [0.000002]0.000040 [0.000000]0.001952 [0.000038]
arrow_table/1000/250/offset-0fresh0.137592 [0.000022]0.005404 [0.000610]0.000111 [0.000002]0.000041 [0.000000]0.143149 [0.000630]
arrow_table/1000/250/offset-0warm0.000559 [0.000010]0.001392 [0.000004]0.000081 [0.000000]0.000040 [0.000001]0.002072 [0.000007]
arrow_table/1000/250/offset-3fresh0.150680 [0.022900]0.005096 [0.000170]0.000108 [0.000004]0.000041 [0.000001]0.155925 [0.023075]
arrow_table/1000/250/offset-3warm0.000605 [0.000031]0.001426 [0.000032]0.000087 [0.000002]0.000041 [0.000001]0.002160 [0.000066]
arrow_table/1000/1000/offset-0fresh0.133647 [0.003753]0.005536 [0.000207]0.000121 [0.000007]0.000041 [0.000001]0.139344 [0.003540]
arrow_table/1000/1000/offset-0warm0.000452 [0.000030]0.001355 [0.000022]0.000078 [0.000001]0.000040 [0.000001]0.001924 [0.000008]
arrow_table/1000/1000/offset-3fresh0.137791 [0.003203]0.005474 [0.000415]0.000109 [0.000003]0.000041 [0.000001]0.143415 [0.003622]
arrow_table/1000/1000/offset-3warm0.000450 [0.000029]0.001369 [0.000048]0.000080 [0.000000]0.000038 [0.000001]0.001938 [0.000078]
pandas/1000/250/offset-0fresh0.003833 [0.000113]0.007438 [0.001073]0.000098 [0.000001]0.000054 [0.000012]0.011423 [0.001175]
pandas/1000/250/offset-0warm0.000938 [0.000061]0.002005 [0.000048]0.000076 [0.000001]0.000039 [0.000001]0.003057 [0.000110]
pandas/1000/250/offset-3fresh0.004045 [0.000086]0.006382 [0.000138]0.000112 [0.000015]0.000043 [0.000001]0.010583 [0.000241]
pandas/1000/250/offset-3warm0.001082 [0.000022]0.002170 [0.000014]0.000077 [0.000001]0.000040 [0.000002]0.003369 [0.000038]
pandas/1000/1000/offset-0fresh0.003638 [0.000089]0.006912 [0.000136]0.000092 [0.000000]0.000040 [0.000000]0.010682 [0.000047]
pandas/1000/1000/offset-0warm0.001138 [0.000168]0.003028 [0.001083]0.000237 [0.000159]0.000100 [0.000045]0.004502 [0.001455]
pandas/1000/1000/offset-3fresh0.003873 [0.000081]0.006744 [0.000008]0.000097 [0.000005]0.000043 [0.000002]0.010758 [0.000092]
pandas/1000/1000/offset-3warm0.000911 [0.000048]0.002019 [0.000049]0.000077 [0.000002]0.000040 [0.000001]0.003047 [0.000096]
polars/1000/250/offset-0fresh0.138789 [0.001956]0.005677 [0.000005]0.000117 [0.000001]0.000041 [0.000001]0.144624 [0.001949]
polars/1000/250/offset-0warm0.000816 [0.000050]0.001530 [0.000010]0.000086 [0.000001]0.000043 [0.000001]0.002475 [0.000038]
polars/1000/250/offset-3fresh0.131215 [0.001698]0.005264 [0.000382]0.000113 [0.000008]0.000041 [0.000001]0.136633 [0.001325]
polars/1000/250/offset-3warm0.000831 [0.000019]0.001462 [0.000030]0.000091 [0.000004]0.000041 [0.000000]0.002425 [0.000046]
polars/1000/1000/offset-0fresh0.141028 [0.000362]0.005354 [0.000213]0.000118 [0.000006]0.000057 [0.000007]0.146557 [0.000151]
polars/1000/1000/offset-0warm0.000655 [0.000033]0.001338 [0.000003]0.000084 [0.000001]0.000040 [0.000000]0.002117 [0.000034]
polars/1000/1000/offset-3fresh0.137571 [0.010698]0.005695 [0.000683]0.000111 [0.000007]0.000045 [0.000003]0.143421 [0.011390]
polars/1000/1000/offset-3warm0.000660 [0.000001]0.001382 [0.000006]0.000082 [0.000005]0.000040 [0.000001]0.002164 [0.000013]
arrow_stream/1000/250/offset-0fresh0.000098 [0.000004]0.141874 [0.004780]0.000118 [0.000013]0.000043 [0.000002]0.142133 [0.004761]
arrow_stream/1000/250/offset-0warm0.000067 [0.000030]0.001755 [0.000071]0.000086 [0.000002]0.000039 [0.000001]0.001946 [0.000105]
arrow_stream/1000/250/offset-3fresh0.000084 [0.000010]0.144676 [0.006776]0.000100 [0.000003]0.000040 [0.000000]0.144900 [0.006790]
arrow_stream/1000/250/offset-3warm0.000063 [0.000026]0.001799 [0.000040]0.000084 [0.000008]0.000039 [0.000001]0.001986 [0.000076]
arrow_stream/1000/1000/offset-0fresh0.000089 [0.000001]0.138067 [0.001035]0.000097 [0.000000]0.000039 [0.000000]0.138292 [0.001036]
arrow_stream/1000/1000/offset-0warm0.000065 [0.000029]0.001704 [0.000050]0.000082 [0.000004]0.000041 [0.000001]0.001891 [0.000084]
arrow_stream/1000/1000/offset-3fresh0.000095 [0.000006]0.139350 [0.005586]0.000096 [0.000002]0.000039 [0.000000]0.139580 [0.005593]
arrow_stream/1000/1000/offset-3warm0.000066 [0.000025]0.001603 [0.000013]0.000088 [0.000008]0.000041 [0.000001]0.001798 [0.000048]
arrow_stream/4000/250/offset-3fresh0.000080 [0.000007]0.145258 [0.004314]0.000107 [0.000001]0.000041 [0.000001]0.145487 [0.004321]
arrow_stream/4000/250/offset-3warm0.000067 [0.000029]0.005836 [0.000030]0.000086 [0.000000]0.000042 [0.000001]0.006031 [0.000059]
02 / Rust profiling

Explore the pipeline.

Repeated in-process measurements, from scan and column profiling to full report assembly. Each case links to its Criterion distributions and estimates.

2 scenario groups shown

Only scenarios measured in this artifact appear here. Routine CI runs the CSV and full-analysis groups; full runs include row counting, scaling and larger inputs. Confidence intervals below describe estimated means.

Observations within the Rust suite
  • csv_parsing climbs from 15.25 MiB/s on tiny inputs to 16.40 MiB/s on medium inputs. Compare the full distributions before attributing the difference.
Read the evidence

A number needs its context.

  • Different measurement boundaries. Criterion measures Rust operations; Python comparisons also cover file reads and, in fresh processes, imports and startup.
  • Variation stays visible. Python shows median and IQR. Criterion shows mean confidence intervals and links to its full distributions.
  • Shared CI is a signal. Runner load and hardware vary. Confirm changes on an idle, controlled host before making performance claims.
  • Scope stays explicit. Tool summaries differ. Input-integrity checks do not establish complete cross-tool metric parity.
Run it yourself

From checkout to evidence.

Use the repository's Rust toolchain and uv. The comparison environment is locked and separate from the dependency-free Python wheel.

# Existing Rust scenarios
cargo bench --bench benchmarks

# Four-tool comparison (from repository root)
uv run --project benches --locked \
  --reinstall-package dataprof python \
  .github/scripts/benchmark_comparison.py
Full protocol, cache controls & repeatability checks ↗