Featured result
Can AI Build a 24/7 Support Schedule? We Tested It
The first schedule broke three rules. We gave the AI one correction pass, then checked every shift with formulas.
- First draft
- 4/10
- One correction
- 10/10
- Verdict
- worked
Auditable AI field tests
Real tasks. Real prompts. Real results. No recycled tool lists, no magic-demo claims, and no failures edited out.
Featured result
The first schedule broke three rules. We gave the AI one correction pass, then checked every shift with formulas.
Every published test includes
Completed field tests
Every score is tied to a visible task, exact prompts, two saved outputs, and five checks. Synthetic benchmarks stay clearly labelled.
Completed field testSynthetic benchmark
The completed LFT-014 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Rubric calibration, while 1 check remained unresolved.
Completed field testSynthetic benchmark
The completed WFT-057 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Claims Triage, while 1 check remained unresolved.
Completed field testSynthetic benchmark
This completed synthetic Incident Evidence field test asked the session to preserve useful evidence after a security alert, preserved an actual five-row incident evidence chain-of-custody ledger, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.
Completed field testSynthetic benchmark
The completed WFT-014 synthetic field test stopped at 6/10: three of five Document Redaction checks passed after one correction, but Document Redaction task fidelity [WFT-014] and Document Redaction handoff usability [WFT-014] remained unsupported.
Completed field testFile-backed test
The first schedule broke three rules. We gave the AI one correction pass, then checked every shift with formulas.
Completed field testSynthetic benchmark
The completed LFT-028 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Nonvisual access, while 0 checks remained unresolved.
Completed field testSynthetic benchmark
The completed WFT-035 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Travel Planning, while 1 check remained unresolved.
Completed field testSynthetic benchmark
This completed synthetic File Cleanup field test asked the session to find duplicate files without deleting originals, preserved an actual five-row duplicate file classification audit, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.
Completed field testSynthetic benchmark
The completed LFT-057 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Rubric Reliability, while 1 check remained unresolved.
Set the task and pass/fail rules before prompting.
Save the exact words, tool, model disclosure, and date.
Check the output independently; do not trust confident prose.
Keep the mistakes, correction, score, and final artifact together.
Two honest evidence modes
File-backed tests preserve inspectable artifacts. Synthetic benchmarks preserve the generated input, Codex multi-agent run, honest model disclosure, prompts, outputs, check evidence, and limitations. Every case says which mode produced its result.
Read the evidence policyTransparency note
WeTriedAI currently runs no ads, analytics, affiliate links, or tracking cookies. If that changes, the relevant pages and disclosures will change first.
Our editorial policy →