Auditable AI field tests

We try AI on real work—then show you exactly what happened.

Real tasks. Real prompts. Real results. No recycled tool lists, no magic-demo claims, and no failures edited out.

Featured field test 001Aug 9, 2026

Featured result

Can AI Build a 24/7 Support Schedule? We Tested It

The first schedule broke three rules. We gave the AI one correction pass, then checked every shift with formulas.

First draft
4/10
One correction
10/10
Verdict
worked

Every published test includes

  • Exact prompt
  • Untouched first result
  • Declared correction limit
  • Five visible checks
  • Downloadable evidence

Completed field tests

The answer is in the record.

Every score is tied to a visible task, exact prompts, two saved outputs, and five checks. Synthetic benchmarks stay clearly labelled.

The WeTriedAI method

One task. One baseline. Evidence before opinion.

Read the complete method →
  1. 01Define

    Set the task and pass/fail rules before prompting.

  2. 02Prompt

    Save the exact words, tool, model disclosure, and date.

  3. 03Verify

    Check the output independently; do not trust confident prose.

  4. 04Publish

    Keep the mistakes, correction, score, and final artifact together.

Two honest evidence modes

Completed does not mean undisclosed.

File-backed tests preserve inspectable artifacts. Synthetic benchmarks preserve the generated input, Codex multi-agent run, honest model disclosure, prompts, outputs, check evidence, and limitations. Every case says which mode produced its result.

Read the evidence policy

Transparency note

WeTriedAI currently runs no ads, analytics, affiliate links, or tracking cookies. If that changes, the relevant pages and disclosures will change first.

Our editorial policy →