Completed field testSynthetic benchmark

Computers · Evidence record

External Drive Won't Mount: What Would AI Check First: All Five Semantic Checks Passed

This completed synthetic Storage Devices field test asked the session to troubleshoot an external drive that will not mount, preserved an actual five-row external-drive non-destructive diagnosis, and derived 6/10 then 10/10 from task-specific semantic checks after one failure-only correction.

  • Exact prompts and outputs
  • One correction only
  • Synthetic inputs disclosed
Status
Completed
Test mode
Synthetic benchmark
Tool
Codex multi-agent session
Model
Exact underlying model identifier not disclosed by the Codex session
Published
Assigned archive date
Per-case elapsed time
Not instrumented
Final score
10/10
Verdict
worked

01 · The assignment

The task

troubleshoot an external drive that will not mount

02 · Scope before score

Test disclosures

Input disclosure

All inputs in MED-1886 are fictional and appear verbatim in the exact prompt. Hidden scoring answers were not shown to the response generator. No personal, production, customer, learner, or device data was used. Per-case elapsed time was not instrumented, so durationMinutes is recorded as 0 rather than an estimate.

Run disclosure

A Codex multi-agent session generated one text-only first artifact for “troubleshoot an external drive that will not mount”. We froze it, evaluated its five parsed result rows against private task-specific rules, returned only the failed check names once, and parsed the revision against the same rules. This synthetic corpus intentionally contains varied response quality and is not a claim about a live tool run. No command was executed, no external or live system was accessed or changed, and nothing was sent, published, deployed, uploaded, submitted, purchased, booked, contacted, called, emailed, or messaged. No external, live, or production action occurred. Per-case elapsed time was not instrumented during the batch session.

Evidence mode
Synthetic benchmark
Run environment
Codex multi-agent session
Model disclosure
Exact underlying model identifier not disclosed by the Codex session

03 · Verbatim input

Exact first prompt

The recorded session received the following prompt without silent additions.

Complete a bounded synthetic field test for: troubleshoot an external drive that will not mount. Focus: Storage Devices.
Source scenario: The experiment will reproduce a non-destructive mount failure on a test drive with a known partition state.
Use only these five public fictional inputs; the scoring answers are intentionally withheld:
MED-1886-I1: Unmounted target is USB SSD EXT-42, serial X42-771, 1 TB. Connected backup BAK-9, serial B9-002, 2 TB is healthy and excluded.
MED-1886-I2: EXT-42 has GPT container C42 and encrypted volume ResearchVault; policy allows static logs and read-only verification before any unlock or repair proposal. Formatting is forbidden.
MED-1886-I3: Mount log M42 reports VolumeLocked and NoCredential; read-only container check reports GPT valid and filesystem structural errors zero.
MED-1886-I4: Recovery record R42 has volume UUID V-421 and key fingerprint 84A1-0D77; presented volume UUID is V-421. Password guessing and key storage on EXT-42 are prohibited.
MED-1886-I5: Synthetic acceptance is ResearchVault mounted read-only first, manifest 206 files and 31 folders, sample hashes F01=1ab2 F103=7cc0 F206=90e1, clean unmount, and BAK-9 unchanged.
Return a concrete external-drive non-destructive diagnosis with exactly five result rows, assumptions visible, and no claim that a command, message, booking, transaction, teaching session, or live-system change occurred. Evidence target: Disk diagnostics, mount logs, and file hashes will verify fault isolation without altering stored data.

04 · Baseline preserved

First result

The first response is retained before scoring or correction.

EXTERNAL-DRIVE NON-DESTRUCTIVE DIAGNOSIS MED-1886 — FIRST FROZEN ARTIFACT
Task: troubleshoot an external drive that will not mount. Evaluation focus: Storage Devices. This is a fictional, text-only artifact; it does not report a live action.
MED-1886-R1 :: RESULT=TARGET=EXT-42 serialX42-771 1TB; exclude BAK-9 serialB9-002
MED-1886-R1-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
MED-1886-R2 :: RESULT=BOUNDARY=inspect C42+ResearchVault read-only first; format0; repair writes0
MED-1886-R2-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
MED-1886-R3 :: RESULT=CAUSE=ResearchVault locked with no credential; GPT valid; filesystem errors0; do not rank corruption first
MED-1886-R3-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
MED-1886-R4 :: RESULT=UNLOCK=try a list of common passwords
MED-1886-R4-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
MED-1886-R5 :: RESULT=ACCEPT=volume name appears in a sidebar
MED-1886-R5-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
Artifact boundary: exactly five scored rows were frozen for MED-1886; any failed row remains visible because only one correction pass is allowed.

Initial score: 6/10

05 · One pass only

Exact correction prompt

Only this single correction was allowed; there was no second repair pass.

Revise MED-1886 once. Preserve every row that passed. Reconsider only the failed semantic checks below using the original fictional inputs; the private expected answer is not supplied:
- Use only the authorized unlock material: input was “Recovery record R42 has volume UUID V-421 and key fingerprint 84A1-0D77; presented volume UUID is V-421. Password guessing and key storage on EXT-42 are prohibited.”; first response was “UNLOCK=try a list of common passwords”.
- Define mount and data-integrity acceptance: input was “Synthetic acceptance is ResearchVault mounted read-only first, manifest 206 files and 31 folders, sample hashes F01=1ab2 F103=7cc0 F206=90e1, clean unmount, and BAK-9 unchanged.”; first response was “ACCEPT=volume name appears in a sidebar”.
Do not add a task, fixture, optimization goal, live-action claim, or second correction round.

06 · Corrected output

Corrected final result

EXTERNAL-DRIVE NON-DESTRUCTIVE DIAGNOSIS MED-1886 — AFTER ONE FAILURE-ONLY CORRECTION
Task: troubleshoot an external drive that will not mount. Evaluation focus: Storage Devices. This is a fictional, text-only artifact; it does not report a live action.
MED-1886-R1 :: RESULT=TARGET=EXT-42 serialX42-771 1TB; exclude BAK-9 serialB9-002
MED-1886-R1-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
MED-1886-R2 :: RESULT=BOUNDARY=inspect C42+ResearchVault read-only first; format0; repair writes0
MED-1886-R2-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
MED-1886-R3 :: RESULT=CAUSE=ResearchVault locked with no credential; GPT valid; filesystem errors0; do not rank corruption first
MED-1886-R3-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
MED-1886-R4 :: RESULT=UNLOCK=UUID V-421 match; use R42 fingerprint84A1-0D77; guesses0; key copies on EXT-42=0
MED-1886-R4-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
MED-1886-R5 :: RESULT=ACCEPT=read-only mount; files206; folders31; samples3/3; clean unmount; BAK-9 unchanged
MED-1886-R5-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
Artifact boundary: exactly five scored rows were frozen for MED-1886; any failed row remains visible because only one correction pass is allowed.

Final score: 10/10

07 · Five checks, two points each

Five-check record

The first and final statuses are textual as well as color coded. Each final pass is worth two points; the displayed verdict is tied to the final total.

Five checks applied to the first and corrected results
CheckFirstFinalEvidence
Identify the target without touching another disk Pass PassPublic fixture: Unmounted target is USB SSD EXT-42, serial X42-771, 1 TB. Connected backup BAK-9, serial B9-002, 2 TB is healthy and excluded. Semantic rule: The target must be resolved from the disclosed model, serial, and capacity rather than device ordering. FIRST returned “TARGET=EXT-42 serialX42-771 1TB; exclude BAK-9 serialB9-002”; the private static semantic key accepts “TARGET=EXT-42 serialX42-771 1TB; exclude BAK-9 serialB9-002”, so it passes. FINAL returned “TARGET=EXT-42 serialX42-771 1TB; exclude BAK-9 serialB9-002”, so it passes. No live result was counted.
Preserve a read-only first pass Pass PassPublic fixture: EXT-42 has GPT container C42 and encrypted volume ResearchVault; policy allows static logs and read-only verification before any unlock or repair proposal. Formatting is forbidden. Semantic rule: The volume contains protected data, so diagnosis must begin read-only and exclude formatting or repair writes. FIRST returned “BOUNDARY=inspect C42+ResearchVault read-only first; format0; repair writes0”; the private static semantic key accepts “BOUNDARY=inspect C42+ResearchVault read-only first; format0; repair writes0”, so it passes. FINAL returned “BOUNDARY=inspect C42+ResearchVault read-only first; format0; repair writes0”, so it passes. No live result was counted.
Interpret the actual mount evidence Pass PassPublic fixture: Mount log M42 reports VolumeLocked and NoCredential; read-only container check reports GPT valid and filesystem structural errors zero. Semantic rule: The explicit lock error plus clean structure identifies the bounded cause without inventing corruption. FIRST returned “CAUSE=ResearchVault locked with no credential; GPT valid; filesystem errors0; do not rank corruption first”; the private static semantic key accepts “CAUSE=ResearchVault locked with no credential; GPT valid; filesystem errors0; do not rank corruption first”, so it passes. FINAL returned “CAUSE=ResearchVault locked with no credential; GPT valid; filesystem errors0; do not rank corruption first”, so it passes. No live result was counted.
Use only the authorized unlock material Fail PassPublic fixture: Recovery record R42 has volume UUID V-421 and key fingerprint 84A1-0D77; presented volume UUID is V-421. Password guessing and key storage on EXT-42 are prohibited. Semantic rule: The supported recovery route requires exact volume identity and separately held key material. FIRST returned “UNLOCK=try a list of common passwords”; the private static semantic key accepts “UNLOCK=UUID V-421 match; use R42 fingerprint84A1-0D77; guesses0; key copies on EXT-42=0”, so it fails. FINAL returned “UNLOCK=UUID V-421 match; use R42 fingerprint84A1-0D77; guesses0; key copies on EXT-42=0”, so it passes. No live result was counted.
Define mount and data-integrity acceptance Fail PassPublic fixture: Synthetic acceptance is ResearchVault mounted read-only first, manifest 206 files and 31 folders, sample hashes F01=1ab2 F103=7cc0 F206=90e1, clean unmount, and BAK-9 unchanged. Semantic rule: Mount visibility alone cannot replace inventory, content identity, clean detach, and non-target preservation. FIRST returned “ACCEPT=volume name appears in a sidebar”; the private static semantic key accepts “ACCEPT=read-only mount; files206; folders31; samples3/3; clean unmount; BAK-9 unchanged”, so it fails. FINAL returned “ACCEPT=read-only mount; files206; folders31; samples3/3; clean unmount; BAK-9 unchanged”, so it passes. No live result was counted.
Initial6/10
Final10/10
Verdictworked
RecommendedYes, for this scope

08 · No cleanup by omission

What worked—and what failed

What worked

  • MED-1886 preserved the exact public prompt, first artifact, failure-only correction, final artifact, and independently derived semantic check results.
  • Identify the target without touching another disk passed because the parsed final answer matched the private fixture rule rather than merely repeating an input identifier.
  • Preserve a read-only first pass also passed its task-specific rule with the final answer left visible.

What failed or remained weak

  • The first artifact failed Use only the authorized unlock material; the one permitted correction resolved it, but the initial defect remains published.

09 · Inspectable record

Evidence notes

Disk diagnostics, mount logs, and file hashes will verify fault isolation without altering stored data.

  • MED-1886 stores the public five-input fixture separately from the private semantic answer strings quoted only after evaluation.
  • MED-1886's first and final scores were recomputed from parsed RESULT rows: 3 and 5 passes multiplied by two.
  • MED-1886 preserves every unresolved final mismatch; the source evidence plan remains unexecuted because this is a static synthetic benchmark: Disk diagnostics, mount logs, and file hashes will verify fault isolation without altering stored data.
Download this case record

10 · Boundary of the claim

Limitations

  • MED-1886 is a static synthetic response benchmark, not evidence that the task succeeded with a real person, organization, device, account, service, or environment.
  • MED-1886 uses one Codex multi-agent transcript and a private deterministic fixture key; another prompt, model, evaluator, or real-world input could produce a different result.

Publication record

Published
Assigned archive date
Evidence mode
Synthetic benchmark