Test independence
Scores come from criteria fixed before the verdict. We do not change a failed result to protect a tool, fit a headline, or make a downloadable artifact look more impressive.
Tools and AI assistance
AI tools may be both the subject of a test and part of our production workflow. Codex currently performs the semantic, evidence, duplication, build, and rendering checks used for editorial QA. We disclose the tested tool, available model identifier, date, correction limit, and material constraints.
No human semantic or editorial review is claimed. A second Codex agent or automated validator is another automated check, not an independent reviewer. The publication record therefore labels this work as a Codex automated editorial pass with no review independence.
The one-year case schedule was created with an AI-assisted workflow and released from a recorded timetable. Build-time validators reject malformed records, invalid dates, missing disclosures, inconsistent scores, and duplicate identities. Every generated record is labelled as a synthetic benchmark produced in a Codex multi-agent session; it is not presented as a live customer deployment.
The first 200 synthetic cases use assigned archive dates from February 9 through August 9, 2026. Those dates organize the initial catalogue; they do not claim that the local-only site was public or that each batch-generated run occurred on that earlier date. Later cases use their recorded automated release timestamps.
The initial catalogue pages identify August 10, 2026 as their publication date and show the assigned archive date separately. Scheduled cases use the time they are released as their publication date. We do not expose an archive, task, or benchmark date to search engines as if it were a page publication date.
Corrections
Material factual or calculation corrections will be dated on the article. We preserve the original test result unless the test itself was invalid; later model changes belong in a new run or update.
Search quality and scale
A timetable does not make a record worth publishing. Each synthetic case must preserve a distinct bounded input, two substantive outputs, one failure-only correction, five task-specific checks, and limitations that match the evidence. We do not publish synonymized keyword variants or create pages whose only purpose is to capture a search query. A record that cannot meet the evidence standard is rejected by validation rather than used to fill a scheduled slot. A structurally complete scheduled record can remain available by direct URL while staying outside the sitemap and search index pending review.
Public release, search approval, promotion approval, and use recommendation are separate decisions. Neither a Markdown file, a downloadable artifact, a publication date, nor an environment setting grants search eligibility. Every searchable URL requires approved automated evidence, semantic, and near-duplicate checks plus a separate, explicit search-publication decision from the site owner. Owner approval is an authorization to expose that exact reviewed content to search; it is not a human semantic review, independent domain review, recommendation, or expert endorsement. Missing or incomplete approval fails closed: the detail page remains outside search-facing aggregations and the sitemap and carries noindex.
Search discovery is automated only after the owner search decision and public URL checks pass. Owned-channel promotion requires a separate approval bound to the exact copy hash, an explicit channel allowlist, a validity window, and a one-run maximum. Search approval alone never authorizes a social post. We do not automate community comments, direct messages, link exchanges, or posts that hide the synthetic evidence mode.
Generated benchmark cases
Items marked Synthetic benchmark are completed, bounded AI exercises using invented inputs. Each clickable detail page identifies the Codex session, states that the model identifier is unavailable, and shows the task, input disclosure, exact first prompt, first output, one correction prompt, corrected output, five checks, score, evidence notes, and limitations. This record supports the stated benchmark result; it does not establish real-world reliability outside the disclosed setup.
The batch session did not instrument elapsed time for each generated case. Those records use a zero duration as an explicit unavailable value and display Not instrumented; we do not convert batch timing into invented per-case minutes.
The 1,105-record collection is a curated synthetic failure-repair corpus, not a representative benchmark sample. The protocol retained a record only when the first response failed at least one check; all 1,105 retained records also have a higher score after the single correction. This selection bias means the corpus cannot estimate ordinary failure rates, general model performance, correction efficacy on an uncontrolled population, or real-world reliability.
Verdicts and recommendations
A verdict is derived from the five published check rows and must match the final score. We retain the first output and permit only the disclosed correction pass. A score or worked verdict is not a use recommendation. A recommendation requires a separate slug-specific recommendation approval; a high-stakes-adjacent task also requires an approved independent domain review. Without those approvals, the public recommendation is shown as not assessed. Mixed and failed outcomes remain available.
Evidence taxonomy
We record the evidence carrier, input origin, execution mode, and run count separately. A workbook or JSON download describes the artifact type, not whether the input was real or whether an external action occurred. Synthetic fixture, permissioned input, public input, sandbox generation, text-only evaluation, live bounded execution, and repeated runs therefore remain distinct disclosures.
Ads, sponsorships, and affiliates
The site currently contains none. If advertising, sponsorship, free access, or affiliate links are introduced, they will be labelled and will not change the scoring rules. A disclosure will appear on every affected page.
Safety and privacy
Tests use fictional, public, or permissioned inputs. We do not publish credentials, private files, personal schedules, student work, or other sensitive data as evidence.