Completed evidence library · Page 4

Completed field tests

Showing cases 7396 of 201. Every case keeps its prompt, first result, correction, checks, score, and evidence limits together.

ComputersRecord date · Jun 13, 2026

Completed field testSynthetic benchmark

Putting a Fragile Legacy Script on AI's Refactoring Bench

This completed synthetic Refactoring field test asked the session to refactor a fragile legacy script, preserved an actual five-row legacy script refactor diff, and derived 0/10 then 2/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 11, 2026

Completed field testSynthetic benchmark

Checking a Container Image for Leaked Secrets and Risky Defaults

This completed synthetic Container Inspection field test asked the session to check a container image for seeded secrets and risky defaults, preserved an actual five-row container secret and defaults audit, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 9, 2026

Completed field testSynthetic benchmark

The Merge-Conflict Resolution Brief for AI

This completed synthetic Version Control field test asked the session to resolve a complex merge conflict correctly, preserved an actual five-row three-way merge resolution record, and derived 2/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 5, 2026

Completed field testSynthetic benchmark

A Safe AI Triage Plan for a Suspicious Email Attachment

This completed synthetic Attachment Safety field test asked the session to triage a suspicious email attachment safely, preserved an actual five-row suspicious attachment static triage record, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 29, 2026

Completed field testSynthetic benchmark

What Does This Kernel Panic Actually Point To

This completed synthetic Kernel Diagnostics field test asked the session to diagnose a kernel panic from a bounded evidence packet, preserved an actual five-row kernel panic causal analysis, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 29, 2026

Completed field testSynthetic benchmark

Could AI Guide a Repeatable Monitor Calibration

This completed synthetic Display Color field test asked the session to guide a basic monitor color calibration, preserved an actual five-row monitor color calibration measurement record, and derived 4/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 24, 2026

Completed field testSynthetic benchmark

From One OS to Another: Porting Shell Automation

This completed synthetic Portability field test asked the session to port a shell automation between operating systems, preserved an actual five-row cross-platform shell portability matrix, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.