llm_eval 0.2.1
llm_eval: ^0.2.1 copied to clipboard
A Dart test harness for LLM evals: assertion checks over model outputs, an optional LLM-as-judge, and cached responses so CI stays deterministic.
0.2.1 #
- Declare the diagram in
pubspec.yamlso pub.dev renders it on the package page. It was already in the repository and the README, but pub.dev shows only what thescreenshots:field points at, so the page opened with prose where the picture should have been.
0.2.0 #
- Add
EvalReport.toJUnitXml(). The exit-code gate added in 0.1.3 turns a build red; this makes the CI system show which cases went red and why. Each eval case becomes a<testcase>, a model error or errored check becomes an<error>, any other non-passing case a<failure>carrying the checks that failed and the model output that failed them, and a flaky case reports how many attempts passed. GitHub Actions, GitLab, Jenkins, CircleCI and Buildkite all read this format. Model output is arbitrary text, so it is XML-escaped and characters XML 1.0 does not permit are dropped rather than emitted; one stray control byte would otherwise make a parser reject the whole report.
0.1.3 #
- Example: use the suite as a CI gate. It now exits non-zero when any case fails, the way you wire it into a build step, instead of only printing a report.
0.1.2 #
- Docs: sharpen the pub.dev description to lead with the value and the terms people search.
0.1.0 #
Initial release.
EvalCase,EvalSuite, andEvalReportwith Markdown and JSON output.- Built-in checks:
contains,notContains,matches,isValidJson,predicate, and LLM-as-judge scoring withCheck.judge. ResponseCacheinterface in the core and a file-backedFileResponseCacheinpackage:llm_eval/io.dart(atomic writes) for deterministic reruns in CI.- Concurrent case execution with stable result order.
- Repeat runs with a flakiness rate.