Topic 03 · Completed field tests

Computers

Completed tests of AI for installing software, writing code, and troubleshooting concrete computer problems.

Completed cases
63
Synthetic benchmarks
63
File-backed tests
0

Published evidence

Completed field tests

Every case is clickable. Synthetic records disclose their generated inputs, Codex multi-agent run, honest model disclosure, checks, and evidence limits.

ComputersRecord date · Aug 2, 2026

Completed field testSynthetic benchmark

Does the Accessibility Tree Match the Visual Interface: The Correction Reached 6/10

This completed synthetic Accessibility Inspection field test asked the session to verify that an interface accessibility tree matches its visual controls, preserved an actual five-row visual-to-accessibility tree audit, and derived 2/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 18, 2026

Completed field testSynthetic benchmark

Planning a Clean OS Reinstall: All Five Semantic Checks Passed

This completed synthetic OS Installation field test asked the session to plan a clean operating system reinstall, preserved an actual five-row clean os reinstall runbook, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 12, 2026

Completed field testSynthetic benchmark

How Should AI Harden a Home Router Without Breaking Devices: One Verified Gap Remained

This completed synthetic Router Hardening field test asked the session to harden a home router while preserving required device connectivity, preserved an actual five-row router hardening and legacy-compatibility matrix, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jul 11, 2026

Completed field testSynthetic benchmark

Debug a Reproducible Memory Leak: The Correction Reached 6/10

This completed synthetic Memory Debugging field test asked the session to debug a reproducible memory leak, preserved an actual five-row software patch and test record, and derived 2/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 11, 2026

Completed field testSynthetic benchmark

Checking a Container Image for Leaked Secrets and Risky Defaults: All Five Semantic Checks Passed

This completed synthetic Container Inspection field test asked the session to check a container image for seeded secrets and risky defaults, preserved an actual five-row container secret and defaults audit, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 9, 2026

Completed field testSynthetic benchmark

Planning a Zero-Downtime Database Migration: The Correction Reached 6/10

This completed synthetic Database Migration field test asked the session to plan a zero-downtime database migration with a safe rollback, preserved an actual five-row zero-downtime migration traffic and rollback plan, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.