Completed evidence library · Page 2

Completed field tests

Showing cases 2548 of 201. Every case keeps its prompt, first result, correction, checks, score, and evidence limits together.

ComputersRecord date · Aug 1, 2026

Completed field testSynthetic benchmark

Can a Language Model Repair This Broken Boot Configuration

This completed synthetic Boot Recovery field test asked the session to repair a broken boot configuration, preserved an actual five-row virtual boot repair decision log, and derived 4/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 31, 2026

Completed field testSynthetic benchmark

An AI Wi-Fi Channel Plan for a Crowded Floor

This completed synthetic Wireless Planning field test asked the session to plan Wi-Fi channels and power levels for a crowded office floor, preserved an actual five-row office wi-fi channel and power plan, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 29, 2026

Completed field testSynthetic benchmark

Bluetooth Audio Gone Wrong: A Troubleshooting Brief for AI

This completed synthetic Bluetooth Audio field test asked the session to troubleshoot distorted bluetooth audio, preserved an actual five-row bluetooth audio fault-isolation log, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

WorkRecord date · Jul 29, 2026

Completed field testSynthetic benchmark

An AI Pick-Path Plan for a Congested Warehouse

The completed WFT-056 synthetic field test stopped at 6/10: three of five Warehouse Routing checks passed after one correction, but Warehouse Routing source traceability [WFT-056] and Warehouse Routing handoff usability [WFT-056] remained unsupported.

WorkRecord date · Jul 23, 2026

Completed field testSynthetic benchmark

Who Decided What? Auditing Meeting Notes with AI

The completed WFT-068 synthetic field test stopped at 6/10: three of five Decision Auditing checks passed after one correction, but Decision Auditing exception handling [WFT-068] and Decision Auditing source traceability [WFT-068] remained unsupported.

ComputersRecord date · Jul 21, 2026

Completed field testSynthetic benchmark

AI and the Privacy-Conscious Home Network Inventory

This completed synthetic Asset Inventory field test asked the session to create a privacy-conscious home network inventory, preserved an actual five-row privacy-minimized home network inventory, and derived 4/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 18, 2026

Completed field testSynthetic benchmark

Planning a Clean OS Reinstall

This completed synthetic OS Installation field test asked the session to plan a clean operating system reinstall, preserved an actual five-row clean os reinstall runbook, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 15, 2026

Completed field testSynthetic benchmark

A Malware Warning, an Inert File, and AI Triage

This completed synthetic Threat Triage field test asked the session to assess a malware warning without running the file, preserved an actual five-row inert-file malware-warning triage, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 12, 2026

Completed field testSynthetic benchmark

How Should AI Harden a Home Router Without Breaking Devices

This completed synthetic Router Hardening field test asked the session to harden a home router while preserving required device connectivity, preserved an actual five-row router hardening and legacy-compatibility matrix, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.