Topic 03 · Page 3

Computers tests

Showing cases 4963 of 63. Every entry links to the full prompts, outputs, checks, and evidence limits.

ComputersRecord date · Mar 23, 2026

Completed field testSynthetic benchmark

Should AI Untangle This Software Dependency Conflict

This completed synthetic Dependencies field test asked the session to resolve a software dependency conflict, preserved an actual five-row dependency resolution lock report, and derived 2/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 20, 2026

Completed field testSynthetic benchmark

Will This API Rate-Limit Plan Hold Under Bursts

This completed synthetic Rate-Limit Testing field test asked the session to design a rate-limit test for bursty API traffic, preserved an actual five-row tiered api rate-limit test matrix, and derived 0/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 18, 2026

Completed field testSynthetic benchmark

Camera and Microphone Permissions: An AI Lockdown Brief

This completed synthetic Device Privacy field test asked the session to lock down camera and microphone access, preserved an actual five-row camera and microphone permission matrix, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 11, 2026

Completed field testSynthetic benchmark

An AI Checklist for a Low-Risk Firmware Update

This completed synthetic Firmware field test asked the session to plan a low-risk firmware update, preserved an actual five-row firmware update safety checklist, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 9, 2026

Completed field testSynthetic benchmark

Planning a Zero-Downtime Database Migration

This completed synthetic Database Migration field test asked the session to plan a zero-downtime database migration with a safe rollback, preserved an actual five-row zero-downtime migration traffic and rollback plan, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 7, 2026

Completed field testSynthetic benchmark

Retired Drive, Verifiable Erasure: An AI Planning Task

This completed synthetic Data Disposal field test asked the session to plan verifiable data removal from a retired drive, preserved an actual five-row retired-drive sanitization decision record, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 4, 2026

Completed field testSynthetic benchmark

Where Should AI Start With Intermittent Wi-Fi Drops

This completed synthetic Wi-Fi field test asked the session to diagnose intermittent wi-fi drops, preserved an actual five-row wi-fi drop diagnostic record, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 27, 2026

Completed field testSynthetic benchmark

Mapping a Project's Software Components

This completed synthetic Supply Chain field test asked the session to inventory a project's software components, preserved an actual five-row software bill of materials audit, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 24, 2026

Completed field testSynthetic benchmark

Build a Cleanly Installable CLI Package

This completed synthetic App Packaging field test asked the session to package a command-line application for clean installation, preserved an actual five-row cli package lifecycle record, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 21, 2026

Completed field testSynthetic benchmark

Does an AI Backup Design Survive a Clean Restore

This completed synthetic Backups field test asked the session to design a backup that can actually be restored, preserved an actual five-row restorable backup design, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 21, 2026

Completed field testSynthetic benchmark

Why Does This Certificate Chain Break on One Client

This completed synthetic TLS Diagnosis field test asked the session to debug a certificate-chain failure that affects one client, preserved an actual five-row tls chain diagnostic record, and derived 2/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 19, 2026

Completed field testSynthetic benchmark

Computer Clock Drift: An AI Diagnosis-and-Correction Brief

This completed synthetic Time Sync field test asked the session to diagnose and correct computer clock drift, preserved an actual five-row performance hypothesis and measurement record, and derived 4/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 16, 2026

Completed field testSynthetic benchmark

Would AI Produce a Reproducible Development Environment

This completed synthetic Dev Environments field test asked the session to build a reproducible development environment, preserved an actual five-row reproducible development environment manifest, and derived 2/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 12, 2026

Completed field testSynthetic benchmark

Using AI to Move a Browser Profile Between Computers

This completed synthetic Profile Migration field test asked the session to migrate a browser profile between computers, preserved an actual five-row browser profile migration manifest, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Feb 9, 2026

Completed field testSynthetic benchmark

Missing Screen-Reader Labels: An AI Repair Brief

This completed synthetic Screen Readers field test asked the session to fix missing screen-reader labels in a desktop app, preserved an actual five-row screen-reader label repair matrix, and derived 0/10 then 2/10 from task-specific semantic checks after one failure-only correction.