Completed evidence library · Page 3

Completed field tests

Showing cases 4972 of 201. Every case keeps its prompt, first result, correction, checks, score, and evidence limits together.

ComputersRecord date · Jul 11, 2026

Completed field testSynthetic benchmark

Debug a Reproducible Memory Leak

This completed synthetic Memory Debugging field test asked the session to debug a reproducible memory leak, preserved an actual five-row software patch and test record, and derived 2/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 7, 2026

Completed field testSynthetic benchmark

When Backup Power Fades: An AI Shutdown-Planning Scenario

This completed synthetic Power Resilience field test asked the session to configure an orderly shutdown during a power failure, preserved an actual five-row ups shutdown policy and simulation matrix, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.

WorkRecord date · Jul 4, 2026

Completed field testSynthetic benchmark

Could AI Build a Complete RFP Compliance Matrix

The completed WFT-038 synthetic field test stopped at 6/10: three of five Proposal Compliance checks passed after one correction, but Proposal Compliance exception handling [WFT-038] and Proposal Compliance source traceability [WFT-038] remained unsupported.

ComputersRecord date · Jul 2, 2026

Completed field testSynthetic benchmark

How Well Might AI Audit Keyboard-Only Navigation

This completed synthetic Keyboard Access field test asked the session to audit desktop keyboard navigation, preserved an actual five-row desktop keyboard navigation trace, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 30, 2026

Completed field testSynthetic benchmark

Unexpected Scheduled Tasks: A Defensive AI Audit

This completed synthetic System Auditing field test asked the session to audit unexpected scheduled tasks defensively, preserved an actual five-row scheduled-task defensive classification, and derived 0/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 26, 2026

Completed field testSynthetic benchmark

Ask AI to Repair a Corrupted Office Document

This completed synthetic File Repair field test asked the session to repair a corrupted office document, preserved an actual five-row corrupted document repair report, and derived 4/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 25, 2026

Completed field testSynthetic benchmark

A Slow SQL Query and an AI Triage Protocol

This completed synthetic Query Performance field test asked the session to triage a slow SQL query using a supplied execution plan, preserved an actual five-row software patch and test record, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 22, 2026

Completed field testSynthetic benchmark

Asking AI to Configure Reliable Log Rotation

This completed synthetic Log Management field test asked the session to configure reliable log rotation, preserved an actual five-row service log rotation specification, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 19, 2026

Completed field testSynthetic benchmark

Least-Privilege Folder Sharing Through AI Guidance

This completed synthetic File Permissions field test asked the session to set least-privilege shared-folder permissions, preserved an actual five-row shared-folder least-privilege matrix, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.