Curated failure-repair corpus · Version 2026-09-25.313

Selected AI failures and one correction, as data.

Download every released record from a deliberately selected synthetic failure-repair corpus. Each record exposes the prompt, retained first failure, one correction, checks, scores, and limitations.

Current release snapshot

313 inspectable records

work
105
learning
110
computers
98
Checks per record
5

Selection boundary

This corpus starts with failure by design.

The complete planned corpus contains 1,105 records. All 1,105 were retained only after the first response failed at least one of five checks, and all 1,105 show a higher check score after the single permitted correction. This is a selection rule, not an estimate of how often an AI fails or improves in ordinary use.

The corpus cannot establish baseline pass rate, general model performance, correction effectiveness on an uncontrolled population, or real-world reliability. The underlying model identifier was unavailable and each record is one synthetic run unless its record says otherwise.

License

Reuse with attribution.

The released synthetic dataset downloads on this page are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0). You may share and adapt them, including commercially, provided you credit WeTriedAI, link to this dataset page and the license, and indicate whether you made changes.

This license applies to the dataset downloads published here. It does not relicense third-party trademarks, linked material, or software in this repository.

What is included

A full audit trail, not a score dump.

Every JSON record contains the fictional input disclosure, exact prompt, untouched first result, one correction limited to detected failures, corrected result, five checks, evidence notes, and limitations.

The CSV is deliberately smaller: it is an index for analysis and links back to the complete public case. No record represents a live customer, private machine, external account, or production action.

Released score distribution

Outcomes stay visible.

2/1014records
4/1048records
6/1079records
8/10118records
10/1054records

Evidence boundary

Synthetic means synthetic.

These are curated response evaluations using invented inputs. They measure whether a selected recorded response passed five bounded checks after at most one correction; they do not establish general model reliability or claim an external task was carried out.

Publication timestamps are separated from assigned archive dates. The latest released record in this dataset was published .

Read the test method and scale and disclosure policy before interpreting the data.