Completed evidence library · Page 5

Completed field tests

Showing cases 97120 of 201. Every case keeps its prompt, first result, correction, checks, score, and evidence limits together.

ComputersRecord date · May 21, 2026

Completed field testSynthetic benchmark

Does AI Reproduce an Environment-Specific Software Bug Reliably

This completed synthetic Bug Reproduction field test asked the session to reproduce an environment-specific software bug, preserved an actual five-row environment-specific bug matrix, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

WorkRecord date · May 19, 2026

Completed field testSynthetic benchmark

Which Refund Cases Need Human Review

The completed WFT-054 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Refund Quality Control, while 1 check remained unresolved.

ComputersRecord date · May 18, 2026

Completed field testSynthetic benchmark

Build a Backup Restore Drill Before Disaster Day

This completed synthetic Restore Testing field test asked the session to build a repeatable backup restore drill for a small service, preserved an actual five-row service restore drill record, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 17, 2026

Completed field testSynthetic benchmark

Could AI Suggest Safer Home-Router Settings

This completed synthetic Router Security field test asked the session to harden a home router safely, preserved an actual five-row home router hardening audit, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 14, 2026

Completed field testSynthetic benchmark

External Drive Won't Mount: What Would AI Check First

This completed synthetic Storage Devices field test asked the session to troubleshoot an external drive that will not mount, preserved an actual five-row external-drive non-destructive diagnosis, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 10, 2026

Completed field testSynthetic benchmark

Check a Software Download's Authenticity With AI and Checksums

This completed synthetic Software Integrity field test asked the session to verify that a software download is authentic, preserved an actual five-row software download authenticity verdict, and derived 0/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 3, 2026

Completed field testSynthetic benchmark

How Much Startup Delay Might AI Remove

This completed synthetic Startup Performance field test asked the session to reduce a computer's startup delay, preserved an actual five-row startup delay prioritization ledger, and derived 2/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 30, 2026

Completed field testSynthetic benchmark

Which Browser Extension Permissions Exceed the Job

This completed synthetic Permission Auditing field test asked the session to audit browser extension permissions against declared functionality, preserved an actual five-row extension feature-to-permission audit, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 29, 2026

Completed field testSynthetic benchmark

What Would AI Make of a Kernel Crash Report

This completed synthetic Crash Analysis field test asked the session to explain a kernel crash report, preserved an actual five-row kernel crash report analysis, and derived 4/10 then 6/10 from task-specific semantic checks after one failure-only correction.