Evidence guide · 39 published records

AI for safer setup and system changes

Review tests of configuration, migration, hardening, installation, backups, permissions, and planned technical changes.

The decision this guide supports

Can AI produce a change plan with prerequisites, least privilege, rollback, and a concrete verification step?

Published records
39
Average first score
2.6/10
Average final score
7.5/10
Average gain
+4.9
Worked / mixed / failed
26 / 6 / 7

Averages use 39 records with a disclosed first score. The set contains 39 synthetic benchmarks and 0 file-backed tests; it is a task collection, not a representative model leaderboard.

What the records compare

Evidence before a recommendation.

Preconditions, permission boundaries, backups, rollback, compatibility, and proof that the final state works.

Across this released set, 7 records failed and 6 remained mixed after the single correction. Those outcomes stay in the guide because a useful decision needs the misses as well as the wins.

  1. Verify prerequisites and compatibility before the first change.
  2. Record a tested rollback or recovery path.
  3. Confirm the final state independently instead of trusting a success message.

Representative evidence

Open the prompts and checks.

The sample deliberately includes different verdicts when available. Every card opens to the full first result, correction, final result, and evidence boundary.

ComputersRecord date · Jun 30, 2026

Completed field testSynthetic benchmark

Unexpected Scheduled Tasks: A Defensive AI Audit

This completed synthetic System Auditing field test asked the session to audit unexpected scheduled tasks defensively, preserved an actual five-row scheduled-task defensive classification, and derived 0/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 7, 2026

Completed field testSynthetic benchmark

When Backup Power Fades: An AI Shutdown-Planning Scenario

This completed synthetic Power Resilience field test asked the session to configure an orderly shutdown during a power failure, preserved an actual five-row ups shutdown policy and simulation matrix, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Aug 8, 2026

Completed field testSynthetic benchmark

Ask AI to Find Duplicate Files—Without Touching the Originals

This completed synthetic File Cleanup field test asked the session to find duplicate files without deleting originals, preserved an actual five-row duplicate file classification audit, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 10, 2026

Completed field testSynthetic benchmark

Check a Software Download's Authenticity With AI and Checksums

This completed synthetic Software Integrity field test asked the session to verify that a software download is authentic, preserved an actual five-row software download authenticity verdict, and derived 0/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 12, 2026

Completed field testSynthetic benchmark

Recover a Corrupt Git Branch with Reversible Steps

This completed synthetic Version-Control Recovery field test asked the session to recover a corrupted Git branch without overwriting healthy history, preserved an actual five-row reversible git branch recovery plan, and derived 0/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 20, 2026

Completed field testSynthetic benchmark

Will This API Rate-Limit Plan Hold Under Bursts

This completed synthetic Rate-Limit Testing field test asked the session to design a rate-limit test for bursty API traffic, preserved an actual five-row tiered api rate-limit test matrix, and derived 0/10 then 8/10 from task-specific semantic checks after one failure-only correction.