Curated failure-repair corpus · Version 2026-09-25.313
Selected AI failures and one correction, as data.
Download every released record from a deliberately selected synthetic failure-repair corpus. Each record exposes the prompt, retained first failure, one correction, checks, scores, and limitations.
Current release snapshot
313 inspectable records
- work
- 105
- learning
- 110
- computers
- 98
- Checks per record
- 5
Selection boundary
This corpus starts with failure by design.
The complete planned corpus contains 1,105 records. All 1,105 were retained only after the first response failed at least one of five checks, and all 1,105 show a higher check score after the single permitted correction. This is a selection rule, not an estimate of how often an AI fails or improves in ordinary use.
The corpus cannot establish baseline pass rate, general model performance, correction effectiveness on an uncontrolled population, or real-world reliability. The underlying model identifier was unavailable and each record is one synthetic run unless its record says otherwise.
License
Reuse with attribution.
The released synthetic dataset downloads on this page are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0). You may share and adapt them, including commercially, provided you credit WeTriedAI, link to this dataset page and the license, and indicate whether you made changes.
This license applies to the dataset downloads published here. It does not relicense third-party trademarks, linked material, or software in this repository.
What is included
A full audit trail, not a score dump.
Every JSON record contains the fictional input disclosure, exact prompt, untouched first result, one correction limited to detected failures, corrected result, five checks, evidence notes, and limitations.
The CSV is deliberately smaller: it is an index for analysis and links back to the complete public case. No record represents a live customer, private machine, external account, or production action.
Released score distribution
Outcomes stay visible.
Evidence boundary
Synthetic means synthetic.
These are curated response evaluations using invented inputs. They measure whether a selected recorded response passed five bounded checks after at most one correction; they do not establish general model reliability or claim an external task was carried out.
Publication timestamps are separated from assigned archive dates. The latest released record in this dataset was published .
Read the test method and scale and disclosure policy before interpreting the data.