Evidence guide · 39 published records
AI for safer setup and system changes
Review tests of configuration, migration, hardening, installation, backups, permissions, and planned technical changes.
The decision this guide supports
Can AI produce a change plan with prerequisites, least privilege, rollback, and a concrete verification step?
- Published records
- 39
- Average first score
- 2.6/10
- Average final score
- 7.5/10
- Average gain
- +4.9
- Worked / mixed / failed
- 26 / 6 / 7
Averages use 39 records with a disclosed first score. The set contains 39 synthetic benchmarks and 0 file-backed tests; it is a task collection, not a representative model leaderboard.
What the records compare
Evidence before a recommendation.
Preconditions, permission boundaries, backups, rollback, compatibility, and proof that the final state works.
Across this released set, 7 records failed and 6 remained mixed after the single correction. Those outcomes stay in the guide because a useful decision needs the misses as well as the wins.
- Verify prerequisites and compatibility before the first change.
- Record a tested rollback or recovery path.
- Confirm the final state independently instead of trusting a success message.
Representative evidence
Open the prompts and checks.
The sample deliberately includes different verdicts when available. Every card opens to the full first result, correction, final result, and evidence boundary.
Completed field testSynthetic benchmark
Unexpected Scheduled Tasks: A Defensive AI Audit
This completed synthetic System Auditing field test asked the session to audit unexpected scheduled tasks defensively, preserved an actual five-row scheduled-task defensive classification, and derived 0/10 then 10/10 from task-specific semantic checks after one failure-only correction.
Completed field testSynthetic benchmark
When Backup Power Fades: An AI Shutdown-Planning Scenario
This completed synthetic Power Resilience field test asked the session to configure an orderly shutdown during a power failure, preserved an actual five-row ups shutdown policy and simulation matrix, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.
Completed field testSynthetic benchmark
Ask AI to Find Duplicate Files—Without Touching the Originals
This completed synthetic File Cleanup field test asked the session to find duplicate files without deleting originals, preserved an actual five-row duplicate file classification audit, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.
Completed field testSynthetic benchmark
Check a Software Download's Authenticity With AI and Checksums
This completed synthetic Software Integrity field test asked the session to verify that a software download is authentic, preserved an actual five-row software download authenticity verdict, and derived 0/10 then 10/10 from task-specific semantic checks after one failure-only correction.
Completed field testSynthetic benchmark
Recover a Corrupt Git Branch with Reversible Steps
This completed synthetic Version-Control Recovery field test asked the session to recover a corrupted Git branch without overwriting healthy history, preserved an actual five-row reversible git branch recovery plan, and derived 0/10 then 8/10 from task-specific semantic checks after one failure-only correction.
Completed field testSynthetic benchmark
Will This API Rate-Limit Plan Hold Under Bursts
This completed synthetic Rate-Limit Testing field test asked the session to design a rate-limit test for bursty API traffic, preserved an actual five-row tiered api rate-limit test matrix, and derived 0/10 then 8/10 from task-specific semantic checks after one failure-only correction.