Test independence
Scores come from criteria fixed before the verdict. We do not change a failed result to protect a tool, fit a headline, or make a downloadable artifact look more impressive.
Tools and AI assistance
AI tools may be both the subject of a test and part of our production workflow. The published evidence is checked separately. We disclose the tested tool, available model identifier, date, correction limit, and material constraints.
The one-year case schedule was created with an AI-assisted workflow and released from a recorded timetable. Build-time validators reject malformed records, invalid dates, missing disclosures, inconsistent scores, and duplicate identities. Every generated record is labelled as a synthetic benchmark produced in a Codex multi-agent session; it is not presented as a live customer deployment.
The first 200 synthetic cases use assigned archive dates from February 9 through August 9, 2026. Those dates organize the initial catalogue; they do not claim that the local-only site was public or that each batch-generated run occurred on that earlier date. Later cases use their recorded automated release timestamps.
The initial catalogue pages identify August 10, 2026 as their publication date and show the assigned archive date separately. Scheduled cases use the time they are released as their publication date. We do not expose an archive, task, or benchmark date to search engines as if it were a page publication date.
Corrections
Material factual or calculation corrections will be dated on the article. We preserve the original test result unless the test itself was invalid; later model changes belong in a new run or update.
Search quality and scale
A timetable does not make a record worth publishing. Each synthetic case must preserve a distinct bounded input, two substantive outputs, one failure-only correction, five task-specific checks, and limitations that match the evidence. We do not publish synonymized keyword variants or create pages whose only purpose is to capture a search query. A record that cannot meet the evidence standard should remain outside public routes and the sitemap rather than fill a scheduled slot.
Generated benchmark cases
Items marked Synthetic benchmark are completed, bounded AI exercises using invented inputs. Each clickable detail page identifies the Codex session, states that the model identifier is unavailable, and shows the task, input disclosure, exact first prompt, first output, one correction prompt, corrected output, five checks, score, evidence notes, and limitations. This record supports the stated benchmark result; it does not establish real-world reliability outside the disclosed setup.
The batch session did not instrument elapsed time for each generated case. Those records use a zero duration as an explicit unavailable value and display Not instrumented; we do not convert batch timing into invented per-case minutes.
Verdicts and recommendations
A verdict is derived from the five published check rows and must match the final score. We retain the first output and permit only the disclosed correction pass. Recommendations apply only to the bounded task and evidence shown; mixed and failed outcomes remain published.
Ads, sponsorships, and affiliates
The site currently contains none. If advertising, sponsorship, free access, or affiliate links are introduced, they will be labelled and will not change the scoring rules. A disclosure will appear on every affected page.
Safety and privacy
Tests use fictional, public, or permissioned inputs. We do not publish credentials, private files, personal schedules, student work, or other sensitive data as evidence.