Run repeatable, local checks over exported language-model tasks. This is a reporter for a deliberately narrow question: did the supplied responses satisfy named literal assertions and declared retry/cost budgets? It does not ask a live model, infer factual truth, or turn a manual judgement into a score.
This walkthrough uses the tool's public README and checked-in example files. Run the command from a repository checkout with Node.js 22+; inspect the source before using it on your own files.
Run the checked-in example
node bin/llm-evaluation-harness.mjs --dataset examples/passing.json
node bin/llm-evaluation-harness.mjs --dataset examples/failing.json --json
node bin/llm-evaluation-harness.mjs --dataset examples/passing.json --responses examples/local-responses.json --json
npm run checkRead the result
The first command exits 0, the second exits 1, and the third uses only the named local response export. stdout is one JSON report and a final newline; without --json a short human summary goes to stderr. --help is on stderr.
Where this check stops
This local walkthrough does not establish the state of a live production system or replace the limits documented in the repository.
Before adapting the command to your own workflow, review the accepted inputs, exit codes and safety boundaries in the README.
Compiled with AI assistance from checked-in public documentation and example scripts. Run the example and review the repository's current documentation before relying on its result.