Evidence guide · 20 published records
AI for records, audits, and reconciliation
See how AI handles discrepancies, logs, inventories, cash records, reports, and other checkable operational data.
The decision this guide supports
Does the result preserve source facts and isolate discrepancies, or merely produce a plausible summary?
- Published records
- 20
- Average first score
- 5/10
- Average final score
- 7.3/10
- Average gain
- +2.3
- Worked / mixed / failed
- 13 / 4 / 3
Averages use 20 records with a disclosed first score. The set contains 20 synthetic benchmarks and 0 file-backed tests; it is a task collection, not a representative model leaderboard.
What the records compare
Evidence before a recommendation.
Source-to-output traceability, arithmetic, missing records, exception handling, and reproducible checks.
Across this released set, 3 records failed and 4 remained mixed after the single correction. Those outcomes stay in the guide because a useful decision needs the misses as well as the wins.
- Keep the source records available for line-by-line verification.
- Use formulas or deterministic rules for arithmetic and completeness.
- Require an exception list instead of allowing silent assumptions.
Representative evidence
Open the prompts and checks.
The sample deliberately includes different verdicts when available. Every card opens to the full first result, correction, final result, and evidence boundary.
Completed field testSynthetic benchmark
Updating a Records Retention Schedule from Policy Amendments
The completed WFT-046 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Records Retention, while 0 checks remained unresolved.
Completed field testSynthetic benchmark
AI Contract Redaction Under a Defined Confidentiality Policy
The completed WFT-014 synthetic field test stopped at 6/10: three of five Document Redaction checks passed after one correction, but Document Redaction task fidelity [WFT-014] and Document Redaction handoff usability [WFT-014] remained unsupported.
Completed field testSynthetic benchmark
AI Procedure Version Comparison with a Paragraph-Level Change Log
The completed WFT-049 synthetic field test finished at 4/10 and was not recommended: only two of five Version Comparison checks passed after the permitted correction.
Completed field testSynthetic benchmark
How Well Can AI Personalize Renewal Outreach from Customer Records
The completed WFT-011 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Renewal Outreach, while 0 checks remained unresolved.
Completed field testSynthetic benchmark
Where Do These Expense Receipts Belong? An AI Reconciliation Protocol
The completed WFT-051 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Expense Reconciliation, while 0 checks remained unresolved.
Completed field testSynthetic benchmark
Could AI Find Duplicate CRM Records Without False Merges
The completed WFT-028 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Record Deduplication, while 1 check remained unresolved.