Evidence guide · 20 published records

AI for records, audits, and reconciliation

See how AI handles discrepancies, logs, inventories, cash records, reports, and other checkable operational data.

The decision this guide supports

Does the result preserve source facts and isolate discrepancies, or merely produce a plausible summary?

Published records
20
Average first score
5/10
Average final score
7.3/10
Average gain
+2.3
Worked / mixed / failed
13 / 4 / 3

Averages use 20 records with a disclosed first score. The set contains 20 synthetic benchmarks and 0 file-backed tests; it is a task collection, not a representative model leaderboard.

What the records compare

Evidence before a recommendation.

Source-to-output traceability, arithmetic, missing records, exception handling, and reproducible checks.

Across this released set, 3 records failed and 4 remained mixed after the single correction. Those outcomes stay in the guide because a useful decision needs the misses as well as the wins.

  1. Keep the source records available for line-by-line verification.
  2. Use formulas or deterministic rules for arithmetic and completeness.
  3. Require an exception list instead of allowing silent assumptions.

Representative evidence

Open the prompts and checks.

The sample deliberately includes different verdicts when available. Every card opens to the full first result, correction, final result, and evidence boundary.