Completed evidence library · Page 8

Completed field tests

Showing cases 169192 of 201. Every case keeps its prompt, first result, correction, checks, score, and evidence limits together.

ComputersRecord date · Mar 4, 2026

Completed field testSynthetic benchmark

Where Should AI Start With Intermittent Wi-Fi Drops

This completed synthetic Wi-Fi field test asked the session to diagnose intermittent wi-fi drops, preserved an actual five-row wi-fi drop diagnostic record, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

LearningRecord date · Feb 28, 2026

Completed field testSynthetic benchmark

A Beginner's Systematic Debugging Lesson with AI

The completed LFT-036 synthetic field test stopped at 6/10: three of five Debugging habits checks passed after one correction, but Debugging habits objective fit [LFT-036] and Debugging habits content accuracy [LFT-036] remained unsupported.

ComputersRecord date · Feb 27, 2026

Completed field testSynthetic benchmark

Mapping a Project's Software Components

This completed synthetic Supply Chain field test asked the session to inventory a project's software components, preserved an actual five-row software bill of materials audit, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 24, 2026

Completed field testSynthetic benchmark

Build a Cleanly Installable CLI Package

This completed synthetic App Packaging field test asked the session to package a command-line application for clean installation, preserved an actual five-row cli package lifecycle record, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 21, 2026

Completed field testSynthetic benchmark

Does an AI Backup Design Survive a Clean Restore

This completed synthetic Backups field test asked the session to design a backup that can actually be restored, preserved an actual five-row restorable backup design, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 21, 2026

Completed field testSynthetic benchmark

Why Does This Certificate Chain Break on One Client

This completed synthetic TLS Diagnosis field test asked the session to debug a certificate-chain failure that affects one client, preserved an actual five-row tls chain diagnostic record, and derived 2/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 19, 2026

Completed field testSynthetic benchmark

Computer Clock Drift: An AI Diagnosis-and-Correction Brief

This completed synthetic Time Sync field test asked the session to diagnose and correct computer clock drift, preserved an actual five-row performance hypothesis and measurement record, and derived 4/10 then 6/10 from task-specific semantic checks after one failure-only correction.

LearningRecord date · Feb 19, 2026

Completed field testSynthetic benchmark

An AI-Led Algebra Diagnostic Interview

The completed LFT-002 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Algebra diagnosis, while 1 check remained unresolved.