Topic 03 · Page 2

Computers tests

Showing cases 2548 of 63. Every entry links to the full prompts, outputs, checks, and evidence limits.

ComputersRecord date · Jun 13, 2026

Completed field testSynthetic benchmark

Putting a Fragile Legacy Script on AI's Refactoring Bench

This completed synthetic Refactoring field test asked the session to refactor a fragile legacy script, preserved an actual five-row legacy script refactor diff, and derived 0/10 then 2/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 11, 2026

Completed field testSynthetic benchmark

Checking a Container Image for Leaked Secrets and Risky Defaults

This completed synthetic Container Inspection field test asked the session to check a container image for seeded secrets and risky defaults, preserved an actual five-row container secret and defaults audit, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 9, 2026

Completed field testSynthetic benchmark

The Merge-Conflict Resolution Brief for AI

This completed synthetic Version Control field test asked the session to resolve a complex merge conflict correctly, preserved an actual five-row three-way merge resolution record, and derived 2/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Jun 5, 2026

Completed field testSynthetic benchmark

A Safe AI Triage Plan for a Suspicious Email Attachment

This completed synthetic Attachment Safety field test asked the session to triage a suspicious email attachment safely, preserved an actual five-row suspicious attachment static triage record, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 29, 2026

Completed field testSynthetic benchmark

What Does This Kernel Panic Actually Point To

This completed synthetic Kernel Diagnostics field test asked the session to diagnose a kernel panic from a bounded evidence packet, preserved an actual five-row kernel panic causal analysis, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 29, 2026

Completed field testSynthetic benchmark

Could AI Guide a Repeatable Monitor Calibration

This completed synthetic Display Color field test asked the session to guide a basic monitor color calibration, preserved an actual five-row monitor color calibration measurement record, and derived 4/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 24, 2026

Completed field testSynthetic benchmark

From One OS to Another: Porting Shell Automation

This completed synthetic Portability field test asked the session to port a shell automation between operating systems, preserved an actual five-row cross-platform shell portability matrix, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 21, 2026

Completed field testSynthetic benchmark

Does AI Reproduce an Environment-Specific Software Bug Reliably

This completed synthetic Bug Reproduction field test asked the session to reproduce an environment-specific software bug, preserved an actual five-row environment-specific bug matrix, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 18, 2026

Completed field testSynthetic benchmark

Build a Backup Restore Drill Before Disaster Day

This completed synthetic Restore Testing field test asked the session to build a repeatable backup restore drill for a small service, preserved an actual five-row service restore drill record, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 17, 2026

Completed field testSynthetic benchmark

Could AI Suggest Safer Home-Router Settings

This completed synthetic Router Security field test asked the session to harden a home router safely, preserved an actual five-row home router hardening audit, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 14, 2026

Completed field testSynthetic benchmark

External Drive Won't Mount: What Would AI Check First

This completed synthetic Storage Devices field test asked the session to troubleshoot an external drive that will not mount, preserved an actual five-row external-drive non-destructive diagnosis, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 10, 2026

Completed field testSynthetic benchmark

Check a Software Download's Authenticity With AI and Checksums

This completed synthetic Software Integrity field test asked the session to verify that a software download is authentic, preserved an actual five-row software download authenticity verdict, and derived 0/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · May 3, 2026

Completed field testSynthetic benchmark

How Much Startup Delay Might AI Remove

This completed synthetic Startup Performance field test asked the session to reduce a computer's startup delay, preserved an actual five-row startup delay prioritization ledger, and derived 2/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 30, 2026

Completed field testSynthetic benchmark

Which Browser Extension Permissions Exceed the Job

This completed synthetic Permission Auditing field test asked the session to audit browser extension permissions against declared functionality, preserved an actual five-row extension feature-to-permission audit, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 29, 2026

Completed field testSynthetic benchmark

What Would AI Make of a Kernel Crash Report

This completed synthetic Crash Analysis field test asked the session to explain a kernel crash report, preserved an actual five-row kernel crash report analysis, and derived 4/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 25, 2026

Completed field testSynthetic benchmark

Synthetic Secrets in Git History: An AI Detection Challenge

This completed synthetic Secret Detection field test asked the session to find exposed secrets in a sample repository, preserved an actual five-row seeded repository secret-detection report, and derived 4/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 19, 2026

Completed field testSynthetic benchmark

Will AI Automate Backup Verification Across Broken Fixtures

This completed synthetic Backup Testing field test asked the session to automate backup verification, preserved an actual five-row read-only backup verification matrix, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 13, 2026

Completed field testSynthetic benchmark

Browser Notification Spam: A Reversible Cleanup Task for AI

This completed synthetic Browser Hygiene field test asked the session to remove unwanted browser notification spam, preserved an actual five-row browser notification permission cleanup, and derived 2/10 then 6/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 12, 2026

Completed field testSynthetic benchmark

Recover a Corrupt Git Branch with Reversible Steps

This completed synthetic Version-Control Recovery field test asked the session to recover a corrupted Git branch without overwriting healthy history, preserved an actual five-row reversible git branch recovery plan, and derived 0/10 then 8/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 10, 2026

Completed field testSynthetic benchmark

Battery Life Without a Broken Workflow: An AI Tuning Brief

This completed synthetic Battery Life field test asked the session to extend laptop battery life without crippling usability, preserved an actual five-row laptop battery usability trial, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 7, 2026

Completed field testSynthetic benchmark

Is an AI-Configured Everyday Account Truly Least Privilege

This completed synthetic Account Security field test asked the session to configure a safer everyday user account, preserved an actual five-row least-privilege account access matrix, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Apr 2, 2026

Completed field testSynthetic benchmark

Ask AI for a Readable High-Contrast Terminal Theme

This completed synthetic Visual Access field test asked the session to design a readable high-contrast terminal theme, preserved an actual five-row high-contrast terminal theme specification, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 30, 2026

Completed field testSynthetic benchmark

Do AI-Guided DNS Filters Respect Family-Safety Boundaries

This completed synthetic DNS Filtering field test asked the session to configure family-safe dns filtering, preserved an actual five-row family dns filter policy, and derived 0/10 then 2/10 from task-specific semantic checks after one failure-only correction.

ComputersRecord date · Mar 27, 2026

Completed field testSynthetic benchmark

Use AI to Sort a Photo Library While Originals Stay Put

This completed synthetic Photo Organization field test asked the session to organize a photo library without moving the originals, preserved an actual five-row read-only photo catalog, and derived 4/10 then 8/10 from task-specific semantic checks after one failure-only correction.