Open source · AI & LLM

LLM Evaluation Harness

Evaluate language model behavior with representative tasks and explicit pass rules.

v0.1.0 · Node.js 22+ · MIT

Browse the public repository · View releases

Run repeatable, local checks over exported language-model tasks. This is a reporter for a deliberately narrow question: did the supplied responses satisfy named literal assertions and declared retry/cost budgets? It does not ask a live model, infer factual truth, or turn a manual judgement into a score.

This walkthrough uses the tool's public README and checked-in example files. Run the command from a repository checkout with Node.js 22+; inspect the source before using it on your own files.

Run the checked-in example

node bin/llm-evaluation-harness.mjs --dataset examples/passing.json
node bin/llm-evaluation-harness.mjs --dataset examples/failing.json --json
node bin/llm-evaluation-harness.mjs --dataset examples/passing.json --responses examples/local-responses.json --json
npm run check

Read the result

The first command exits 0, the second exits 1, and the third uses only the named local response export. stdout is one JSON report and a final newline; without --json a short human summary goes to stderr. --help is on stderr.

Where this check stops

This local walkthrough does not establish the state of a live production system or replace the limits documented in the repository.

Before adapting the command to your own workflow, review the accepted inputs, exit codes and safety boundaries in the README.

Compiled with AI assistance from checked-in public documentation and example scripts. Run the example and review the repository's current documentation before relying on its result.