Completed evidence library · Page 7

Completed field tests

Showing cases 145168 of 201. Every case keeps its prompt, first result, correction, checks, score, and evidence limits together.

ComputersRecord date · Mar 30, 2026

Completed field testSynthetic benchmark

Do AI-Guided DNS Filters Respect Family-Safety Boundaries

This completed synthetic DNS Filtering field test asked the session to configure family-safe dns filtering, preserved an actual five-row family dns filter policy, and derived 0/10 then 2/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 27, 2026

Completed field testSynthetic benchmark

Use AI to Sort a Photo Library While Originals Stay Put

This completed synthetic Photo Organization field test asked the session to organize a photo library without moving the originals, preserved an actual five-row read-only photo catalog, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 23, 2026

Completed field testSynthetic benchmark

Should AI Untangle This Software Dependency Conflict

This completed synthetic Dependencies field test asked the session to resolve a software dependency conflict, preserved an actual five-row dependency resolution lock report, and derived 2/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 20, 2026

Completed field testSynthetic benchmark

Will This API Rate-Limit Plan Hold Under Bursts

This completed synthetic Rate-Limit Testing field test asked the session to design a rate-limit test for bursty API traffic, preserved an actual five-row tiered api rate-limit test matrix, and derived 0/10 then 8/10 from task-specific semantic checks after one failure-only correction.

WorkRecord date · Mar 19, 2026

Completed field testSynthetic benchmark

Building an Incident Timeline from Conflicting Records

The completed WFT-026 synthetic field test stopped at 6/10: three of five Incident Review checks passed after one correction, but Incident Review source traceability [WFT-026] and Incident Review handoff usability [WFT-026] remained unsupported.

ComputersRecord date · Mar 18, 2026

Completed field testSynthetic benchmark

Camera and Microphone Permissions: An AI Lockdown Brief

This completed synthetic Device Privacy field test asked the session to lock down camera and microphone access, preserved an actual five-row camera and microphone permission matrix, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

LearningRecord date · Mar 18, 2026

Completed field testSynthetic benchmark

Plan a Museum Observation Guide with AI

The completed LFT-046 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Museum learning, while 1 check remained unresolved.

ComputersRecord date · Mar 11, 2026

Completed field testSynthetic benchmark

An AI Checklist for a Low-Risk Firmware Update

This completed synthetic Firmware field test asked the session to plan a low-risk firmware update, preserved an actual five-row firmware update safety checklist, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 9, 2026

Completed field testSynthetic benchmark

Planning a Zero-Downtime Database Migration

This completed synthetic Database Migration field test asked the session to plan a zero-downtime database migration with a safe rollback, preserved an actual five-row zero-downtime migration traffic and rollback plan, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 7, 2026

Completed field testSynthetic benchmark

Retired Drive, Verifiable Erasure: An AI Planning Task

This completed synthetic Data Disposal field test asked the session to plan verifiable data removal from a retired drive, preserved an actual five-row retired-drive sanitization decision record, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.