Why this site exists
AI demos are usually clean. Real tasks are not. Schedules have availability rules, installations fail, course material needs fact-checking, and tutoring can become answer-giving. We publish the messy middle between a prompt and a usable result.
What we publish
Published tests start with a specific task and end with evidence: a file, screenshot, command log, or other result that can be checked. We include prompts and failures because they are part of the product experience. Our generated cases are clearly marked synthetic benchmarks and disclose the Codex session, unavailable model identifier, invented inputs, and execution limits behind the result.
What the growing library contains
The completed test library spans Work, Learning, and Computers. Every generated case opens to the task, input and run disclosures, exact prompts and results, a 5 × 2 check table, one-correction record, evidence notes, verdict, and limitations. Two or three dated cases are scheduled between 08:00 and 18:00 Shanghai time each day; only cases whose recorded release time has arrived appear on the site.
Who writes the tests
Articles are currently published under the brand byline WeTriedAI Editorial Team. We do not attach invented biographies, credentials, or first-person identities to that byline.
Start with the evidence
Read our testing method, browse the published field tests, or start with the Work, Learning, and Computers collections.