| Benchmark | Mean | Confidence interval | Throughput |
|---|---|---|---|
| parse / tiny | 0.41 ms | 95%: 0.40 ms–0.41 ms | 15.25 MiB/s |
| parse / small | 3.86 ms | 95%: 3.84 ms–3.87 ms | 16.66 MiB/s |
| parse / medium | 41.42 ms | 95%: 40.64 ms–42.42 ms | 16.40 MiB/s |
See how dataprof performs.
From a CSV scan to a complete profile. Explore measured workloads, inspect their variation, and run the same experiments on your own data infrastructure.
One input. Explicit workloads.
File-to-summary timings for 1,000 rows of numeric, text and null data. Every sample checks row counts, column order and nulls.
Warm operations
Repeated CSV reads and fresh summaries after warmup. Adapter import/setup excluded.
Fresh-process operations
A fresh process per sample. Startup, imports, operation and exit included.
How to read these numbers
Diagnostic evidence — no established baseline. Publication mode increases the experiment budget; it does not certify precision. Warm samples share a process within each block. Separate blocks use fresh processes, but share host and OS caches. Within-run intervals do not demonstrate across-run stability; repeat-run IQR overlap is a diagnostic, not a significance test.
Optional resource collection enabled. Energy and peak RSS are per worker, including imports and warmups; warm values cover the entire block, not a single operation. Host energy zones include background load and collector overhead and are never summed. Unavailable counters remain unavailable. See the resource median/IQR tables and raw readings and paired idle baselines. Shared-runner results are collection smoke evidence, not energy-efficiency claims.
Ordered samples and timing boundaries
Seconds in execution order; no first sample is discarded. Process time includes interpreter startup, harness imports, adapter import/setup, operations, validation, garbage collection, IPC and exit. Import/setup is timed directly around the adapter; deferred imports remain inside operations. First operation includes initialization and, for warm workers, is the first warmup. Warmup samples are retained in the raw JSON.
First adapter import after environment preparation (preflight): dataprof: 0.009436; pandas: 0.198340; polars: 0.057147; ydata-profiling: 1.690976. Earlier installs or workflow preflights may have warmed library pages; this is not disk-cold.
| Order | Block | Tool | Mode | Invocation | Process | Import/setup | First operation | Measured operations |
|---|---|---|---|---|---|---|---|---|
| 1 | 1 | polars | cold | first_fixture_operation_after_preflight | 0.161429 | 0.057506 | 0.003090 | 0.003090 |
| 2 | 1 | pandas | cold | first_fixture_operation_after_preflight | 0.387540 | 0.216897 | 0.005493 | 0.005493 |
| 3 | 1 | dataprof | cold | first_fixture_operation_after_preflight | 0.079603 | 0.009814 | 0.004957 | 0.004957 |
| 4 | 1 | ydata-profiling | cold | first_fixture_operation_after_preflight | 2.801696 | 1.679654 | 0.373061 | 0.373061 |
| 5 | 1 | dataprof | cold | subsequent_fresh_process | 0.078546 | 0.009416 | 0.004908 | 0.004908 |
| 6 | 1 | pandas | cold | subsequent_fresh_process | 0.359684 | 0.207151 | 0.005174 | 0.005174 |
| 7 | 1 | ydata-profiling | cold | subsequent_fresh_process | 2.612088 | 1.649741 | 0.206973 | 0.206973 |
| 8 | 1 | polars | cold | subsequent_fresh_process | 0.154886 | 0.059341 | 0.003109 | 0.003109 |
| 9 | 1 | polars | cold | subsequent_fresh_process | 0.153101 | 0.057396 | 0.003060 | 0.003060 |
| 10 | 1 | dataprof | cold | subsequent_fresh_process | 0.074195 | 0.009102 | 0.004786 | 0.004786 |
| 11 | 1 | pandas | cold | subsequent_fresh_process | 0.359792 | 0.208840 | 0.004943 | 0.004943 |
| 12 | 1 | ydata-profiling | cold | subsequent_fresh_process | 2.549270 | 1.582732 | 0.211612 | 0.211612 |
| 13 | 1 | dataprof | warm | subsequent_fresh_process | 0.093122 | 0.009868 | 0.004895 | 0.001664, 0.001682, 0.001583 |
| 14 | 1 | pandas | warm | subsequent_fresh_process | 0.398073 | 0.207412 | 0.005165 | 0.003541, 0.003375, 0.003333 |
| 15 | 1 | ydata-profiling | warm | subsequent_fresh_process | 3.644522 | 1.562768 | 0.317927 | 0.211409, 0.216925, 0.239258 |
| 16 | 1 | polars | warm | subsequent_fresh_process | 0.159008 | 0.056975 | 0.002910 | 0.001042, 0.000968, 0.000916 |
Minimal-worker controls, before each round (seconds): block 1, round 1: 0.017430; block 1, round 2: 0.017411; block 1, round 3: 0.016300.
Control scope: interpreter, minimal JSON worker, IPC and exit, without harness or tool imports. It is context for combined startup costs, not a subtractable import estimate.
Exact workloads and timing table
| Tool | Cold median [IQR] | Warm median [IQR] | Workload |
|---|---|---|---|
| dataprof | 0.078546 [0.002704] | 0.001664 [0.000049] | profile(engine="auto", metrics=["schema", "statistics"]) |
| pandas | 0.359792 [0.013928] | 0.003375 [0.000104] | read_csv + describe(include="all") + null counts |
| polars | 0.154886 [0.004164] | 0.000968 [0.000063] | read_csv + describe + null counts |
| ydata-profiling | 2.612088 [0.126213] | 0.216925 [0.013924] | pandas.read_csv + ProfileReport(minimal=True).description_set |
Machine, versions and fixture identity
- Measured at
- 2026-10-08T07:57:22.883260+00:00
- Operating system
- Linux-6.17.0-1022-azure-x86_64-with-glibc2.39
- Processor
- AMD EPYC 9V45 96-Core Processor
- Samples / warmups per block
- 3 / 1
- Process blocks
- 1
- Experiment mode
- routine
- Declared host controls
- GitHub-hosted runner; resource collection smoke test only
- Library cache
- not evicted; environment preparation and preflight may warm library pages and caches
- Import preflight
- imports only, in disposable workers before measurements; no fixture operations; may warm OS library pages and on-disk library caches, which are not evicted; fresh-process samples are subsequent to preflight, not first host invocations
- Thread request
- 1
- Cold file cache
- warm (residency / eviction unverified)
- Fixture SHA-256
- 1d766b0c21f3a89d78c31e76c9196471a2b5c7a87e6c39bdcff87e44b8e42272
- Checkout
- 85c0850eb934a9f1277231a38129626297e0991c
- Working tree
- clean
- Installed tools
- dataprof 0.12.0, pandas 2.3.3, polars 1.44.2, ydata-profiling 4.18.4
Python / Arrow boundaries
What it costs to hand pandas, polars and Arrow data to dataprof, stage by stage.
How to read these numbers
Diagnostic evidence, no established baseline. Exact serialized column metrics, absence, integers, nulls and order are checked before comparison. Consumer import and profiling share one timer; lazy stream construction happens inside it. End-to-end includes both independent exports. Fresh means a new process, not cold storage.
Peak RSS includes native buffers, imports and warmups across the worker lifetime. Arrow pool observations are partial allocator evidence, not profiler-only memory.
Stage times: median [IQR], seconds
| Producer / rows / chunk / offset | Condition | Prepare | Import + profile | Dict | JSON | End-to-end |
|---|---|---|---|---|---|---|
| arrow_array/1000/250/offset-0 | Skipped: C Array exports one batch; chunked input uses Table or C Stream | |||||
| arrow_array/1000/250/offset-3 | Skipped: C Array exports one batch; chunked input uses Table or C Stream | |||||
| arrow_array/1000/1000/offset-0 | fresh | 0.135536 [0.003819] | 0.005084 [0.000562] | 0.000118 [0.000009] | 0.000053 [0.000014] | 0.140791 [0.003280] |
| arrow_array/1000/1000/offset-0 | warm | 0.000492 [0.000003] | 0.001317 [0.000018] | 0.000085 [0.000001] | 0.000040 [0.000000] | 0.001934 [0.000022] |
| arrow_array/1000/1000/offset-3 | fresh | 0.136664 [0.000184] | 0.004533 [0.000049] | 0.000123 [0.000006] | 0.000041 [0.000001] | 0.141361 [0.000240] |
| arrow_array/1000/1000/offset-3 | warm | 0.000485 [0.000016] | 0.001346 [0.000024] | 0.000081 [0.000002] | 0.000040 [0.000000] | 0.001952 [0.000038] |
| arrow_table/1000/250/offset-0 | fresh | 0.137592 [0.000022] | 0.005404 [0.000610] | 0.000111 [0.000002] | 0.000041 [0.000000] | 0.143149 [0.000630] |
| arrow_table/1000/250/offset-0 | warm | 0.000559 [0.000010] | 0.001392 [0.000004] | 0.000081 [0.000000] | 0.000040 [0.000001] | 0.002072 [0.000007] |
| arrow_table/1000/250/offset-3 | fresh | 0.150680 [0.022900] | 0.005096 [0.000170] | 0.000108 [0.000004] | 0.000041 [0.000001] | 0.155925 [0.023075] |
| arrow_table/1000/250/offset-3 | warm | 0.000605 [0.000031] | 0.001426 [0.000032] | 0.000087 [0.000002] | 0.000041 [0.000001] | 0.002160 [0.000066] |
| arrow_table/1000/1000/offset-0 | fresh | 0.133647 [0.003753] | 0.005536 [0.000207] | 0.000121 [0.000007] | 0.000041 [0.000001] | 0.139344 [0.003540] |
| arrow_table/1000/1000/offset-0 | warm | 0.000452 [0.000030] | 0.001355 [0.000022] | 0.000078 [0.000001] | 0.000040 [0.000001] | 0.001924 [0.000008] |
| arrow_table/1000/1000/offset-3 | fresh | 0.137791 [0.003203] | 0.005474 [0.000415] | 0.000109 [0.000003] | 0.000041 [0.000001] | 0.143415 [0.003622] |
| arrow_table/1000/1000/offset-3 | warm | 0.000450 [0.000029] | 0.001369 [0.000048] | 0.000080 [0.000000] | 0.000038 [0.000001] | 0.001938 [0.000078] |
| pandas/1000/250/offset-0 | fresh | 0.003833 [0.000113] | 0.007438 [0.001073] | 0.000098 [0.000001] | 0.000054 [0.000012] | 0.011423 [0.001175] |
| pandas/1000/250/offset-0 | warm | 0.000938 [0.000061] | 0.002005 [0.000048] | 0.000076 [0.000001] | 0.000039 [0.000001] | 0.003057 [0.000110] |
| pandas/1000/250/offset-3 | fresh | 0.004045 [0.000086] | 0.006382 [0.000138] | 0.000112 [0.000015] | 0.000043 [0.000001] | 0.010583 [0.000241] |
| pandas/1000/250/offset-3 | warm | 0.001082 [0.000022] | 0.002170 [0.000014] | 0.000077 [0.000001] | 0.000040 [0.000002] | 0.003369 [0.000038] |
| pandas/1000/1000/offset-0 | fresh | 0.003638 [0.000089] | 0.006912 [0.000136] | 0.000092 [0.000000] | 0.000040 [0.000000] | 0.010682 [0.000047] |
| pandas/1000/1000/offset-0 | warm | 0.001138 [0.000168] | 0.003028 [0.001083] | 0.000237 [0.000159] | 0.000100 [0.000045] | 0.004502 [0.001455] |
| pandas/1000/1000/offset-3 | fresh | 0.003873 [0.000081] | 0.006744 [0.000008] | 0.000097 [0.000005] | 0.000043 [0.000002] | 0.010758 [0.000092] |
| pandas/1000/1000/offset-3 | warm | 0.000911 [0.000048] | 0.002019 [0.000049] | 0.000077 [0.000002] | 0.000040 [0.000001] | 0.003047 [0.000096] |
| polars/1000/250/offset-0 | fresh | 0.138789 [0.001956] | 0.005677 [0.000005] | 0.000117 [0.000001] | 0.000041 [0.000001] | 0.144624 [0.001949] |
| polars/1000/250/offset-0 | warm | 0.000816 [0.000050] | 0.001530 [0.000010] | 0.000086 [0.000001] | 0.000043 [0.000001] | 0.002475 [0.000038] |
| polars/1000/250/offset-3 | fresh | 0.131215 [0.001698] | 0.005264 [0.000382] | 0.000113 [0.000008] | 0.000041 [0.000001] | 0.136633 [0.001325] |
| polars/1000/250/offset-3 | warm | 0.000831 [0.000019] | 0.001462 [0.000030] | 0.000091 [0.000004] | 0.000041 [0.000000] | 0.002425 [0.000046] |
| polars/1000/1000/offset-0 | fresh | 0.141028 [0.000362] | 0.005354 [0.000213] | 0.000118 [0.000006] | 0.000057 [0.000007] | 0.146557 [0.000151] |
| polars/1000/1000/offset-0 | warm | 0.000655 [0.000033] | 0.001338 [0.000003] | 0.000084 [0.000001] | 0.000040 [0.000000] | 0.002117 [0.000034] |
| polars/1000/1000/offset-3 | fresh | 0.137571 [0.010698] | 0.005695 [0.000683] | 0.000111 [0.000007] | 0.000045 [0.000003] | 0.143421 [0.011390] |
| polars/1000/1000/offset-3 | warm | 0.000660 [0.000001] | 0.001382 [0.000006] | 0.000082 [0.000005] | 0.000040 [0.000001] | 0.002164 [0.000013] |
| arrow_stream/1000/250/offset-0 | fresh | 0.000098 [0.000004] | 0.141874 [0.004780] | 0.000118 [0.000013] | 0.000043 [0.000002] | 0.142133 [0.004761] |
| arrow_stream/1000/250/offset-0 | warm | 0.000067 [0.000030] | 0.001755 [0.000071] | 0.000086 [0.000002] | 0.000039 [0.000001] | 0.001946 [0.000105] |
| arrow_stream/1000/250/offset-3 | fresh | 0.000084 [0.000010] | 0.144676 [0.006776] | 0.000100 [0.000003] | 0.000040 [0.000000] | 0.144900 [0.006790] |
| arrow_stream/1000/250/offset-3 | warm | 0.000063 [0.000026] | 0.001799 [0.000040] | 0.000084 [0.000008] | 0.000039 [0.000001] | 0.001986 [0.000076] |
| arrow_stream/1000/1000/offset-0 | fresh | 0.000089 [0.000001] | 0.138067 [0.001035] | 0.000097 [0.000000] | 0.000039 [0.000000] | 0.138292 [0.001036] |
| arrow_stream/1000/1000/offset-0 | warm | 0.000065 [0.000029] | 0.001704 [0.000050] | 0.000082 [0.000004] | 0.000041 [0.000001] | 0.001891 [0.000084] |
| arrow_stream/1000/1000/offset-3 | fresh | 0.000095 [0.000006] | 0.139350 [0.005586] | 0.000096 [0.000002] | 0.000039 [0.000000] | 0.139580 [0.005593] |
| arrow_stream/1000/1000/offset-3 | warm | 0.000066 [0.000025] | 0.001603 [0.000013] | 0.000088 [0.000008] | 0.000041 [0.000001] | 0.001798 [0.000048] |
| arrow_stream/4000/250/offset-3 | fresh | 0.000080 [0.000007] | 0.145258 [0.004314] | 0.000107 [0.000001] | 0.000041 [0.000001] | 0.145487 [0.004321] |
| arrow_stream/4000/250/offset-3 | warm | 0.000067 [0.000029] | 0.005836 [0.000030] | 0.000086 [0.000000] | 0.000042 [0.000001] | 0.006031 [0.000059] |
Explore the pipeline.
Repeated in-process measurements, from scan and column profiling to full report assembly. Each case links to its Criterion distributions and estimates.
2 scenario groups shown
Only scenarios measured in this artifact appear here. Routine CI runs the CSV and full-analysis groups; full runs include row counting, scaling and larger inputs. Confidence intervals below describe estimated means.
| Benchmark | Mean | Confidence interval | Throughput |
|---|---|---|---|
| analyze / tiny | 0.83 ms | 95%: 0.82 ms–0.83 ms | 7.49 MiB/s |
| analyze / small | 6.69 ms | 95%: 6.65 ms–6.73 ms | 9.61 MiB/s |
| analyze / medium | 62.28 ms | 95%: 61.88 ms–62.92 ms | 10.91 MiB/s |
Observations within the Rust suite
- csv_parsing climbs from 15.25 MiB/s on tiny inputs to 16.40 MiB/s on medium inputs. Compare the full distributions before attributing the difference.
A number needs its context.
- Different measurement boundaries. Criterion measures Rust operations; Python comparisons also cover file reads and, in fresh processes, imports and startup.
- Variation stays visible. Python shows median and IQR. Criterion shows mean confidence intervals and links to its full distributions.
- Shared CI is a signal. Runner load and hardware vary. Confirm changes on an idle, controlled host before making performance claims.
- Scope stays explicit. Tool summaries differ. Input-integrity checks do not establish complete cross-tool metric parity.
From checkout to evidence.
Use the repository's Rust toolchain and uv. The comparison environment is locked and separate from the dependency-free Python wheel.
# Existing Rust scenarios
cargo bench --bench benchmarks
# Four-tool comparison (from repository root)
uv run --project benches --locked \
--reinstall-package dataprof python \
.github/scripts/benchmark_comparison.py
Full protocol, cache controls & repeatability checks ↗