Completed evidence library · Page 6

Completed field tests

Showing cases 121144 of 201. Every case keeps its prompt, first result, correction, checks, score, and evidence limits together.

ComputersRecord date · Apr 25, 2026

Completed field testSynthetic benchmark

Synthetic Secrets in Git History: An AI Detection Challenge

This completed synthetic Secret Detection field test asked the session to find exposed secrets in a sample repository, preserved an actual five-row seeded repository secret-detection report, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

LearningRecord date · Apr 24, 2026

Completed field testSynthetic benchmark

A Faded-Example Plan for Learners with Dyscalculia

The completed LFT-066 synthetic field test stopped at 6/10: three of five Accessible Mathematics checks passed after one correction, but Accessible Mathematics objective fit [LFT-066] and Accessible Mathematics content accuracy [LFT-066] remained unsupported.

ComputersRecord date · Apr 19, 2026

Completed field testSynthetic benchmark

Will AI Automate Backup Verification Across Broken Fixtures

This completed synthetic Backup Testing field test asked the session to automate backup verification, preserved an actual five-row read-only backup verification matrix, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

LearningRecord date · Apr 14, 2026

Completed field testSynthetic benchmark

Map Debate Evidence on Both Sides with AI

The completed LFT-048 synthetic field test stopped at 6/10: three of five Debate reasoning checks passed after one correction, but Debate reasoning objective fit [LFT-048] and Debate reasoning safety and access [LFT-048] remained unsupported.

ComputersRecord date · Apr 13, 2026

Completed field testSynthetic benchmark

Browser Notification Spam: A Reversible Cleanup Task for AI

This completed synthetic Browser Hygiene field test asked the session to remove unwanted browser notification spam, preserved an actual five-row browser notification permission cleanup, and derived 2/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 12, 2026

Completed field testSynthetic benchmark

Recover a Corrupt Git Branch with Reversible Steps

This completed synthetic Version-Control Recovery field test asked the session to recover a corrupted Git branch without overwriting healthy history, preserved an actual five-row reversible git branch recovery plan, and derived 0/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 10, 2026

Completed field testSynthetic benchmark

Battery Life Without a Broken Workflow: An AI Tuning Brief

This completed synthetic Battery Life field test asked the session to extend laptop battery life without crippling usability, preserved an actual five-row laptop battery usability trial, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

LearningRecord date · Apr 9, 2026

Completed field testSynthetic benchmark

Giving Debugging Hints Without Writing the Student's Code

The completed LFT-060 synthetic field test stopped at 6/10: three of five Programming Tutoring checks passed after one correction, but Programming Tutoring evidence traceability [LFT-060] and Programming Tutoring safety and access [LFT-060] remained unsupported.

ComputersRecord date · Apr 7, 2026

Completed field testSynthetic benchmark

Is an AI-Configured Everyday Account Truly Least Privilege

This completed synthetic Account Security field test asked the session to configure a safer everyday user account, preserved an actual five-row least-privilege account access matrix, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 2, 2026

Completed field testSynthetic benchmark

Ask AI for a Readable High-Contrast Terminal Theme

This completed synthetic Visual Access field test asked the session to design a readable high-contrast terminal theme, preserved an actual five-row high-contrast terminal theme specification, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.