A profiler is easy to write and easy to write dishonestly. The dishonest version computes a median over whatever parsed, places a fence around it, finds nothing outside, and prints a clean column. It has told you nothing and made it look like something, and the next person reads the green output as evidence.
This walkthrough uses the tool's public README and checked-in example files. Run the command from a repository checkout with Node.js 22+; inspect the source before using it on your own files.
Run the checked-in example
# A file that matches its baseline: every column gets the verdict it was asked
# for and none of them fails a threshold.
node bin/csv-anomaly-profiler.mjs \
--csv examples/clean/orders.csv \
--baseline examples/clean/baseline.json
# exit 0, status "pass"
# The same shape with a planted outlier, a category the baseline does not list,
# and a column that started going missing.
node bin/csv-anomaly-profiler.mjs \
--csv examples/anomalous/orders.csv \
--baseline examples/anomalous/baseline.json
# exit 1, status "fail"
# Six rows and a column that is part letters: neither supports a verdict.
node bin/csv-anomaly-profiler.mjs --csv examples/incomplete/readings.csv
# exit 2, status "incomplete"Read the result
stdout carries the JSON report and nothing else, so it pipes straight into a parser. The human summary goes to stderr, and --json silences it.
Where this check stops
Measured rather than asserted: the largest configuration these bounds permit -- maxRows 1953, maxColumns 1024, maxDistinctCategories 4096, maxCategoryLength 8 -- over an 18 MB file of 1999872 cells, every value distinct and every column tracked against a baseline, used 398 MB of peak resident memory and 42 s of CPU on one developer machine, and produced a 6.6 MB report. The figures are what one run on one machine did, not a promise; the bounds are the promise.
Before adapting the command to your own workflow, review the accepted inputs, exit codes and safety boundaries in the README.
Compiled with AI assistance from checked-in public documentation and example scripts. Run the example and review the repository's current documentation before relying on its result.