Completed field testSynthetic benchmark

Computers · Evidence record

Why Is the Test Computer Running at High CPU? An AI Diagnosis: The Correction Reached 6/10

This completed synthetic CPU Performance field test asked the session to diagnose unexplained high cpu usage, preserved an actual five-row high-cpu causal trace, and derived 0/10 then 6/10 from task-specific semantic checks after one failure-only correction.

  • Exact prompts and outputs
  • One correction only
  • Synthetic inputs disclosed
Status
Completed
Test mode
Synthetic benchmark
Tool
Codex multi-agent session
Model
Exact underlying model identifier not disclosed by the Codex session
Published
Assigned archive date
Per-case elapsed time
Not instrumented
Final score
6/10
Verdict
mixed

01 · The assignment

The task

diagnose unexplained high cpu usage

02 · Scope before score

Test disclosures

Input disclosure

All inputs in DHC-4608 are fictional and appear verbatim in the exact prompt. Hidden scoring answers were not shown to the response generator. No personal, production, customer, learner, or device data was used. Per-case elapsed time was not instrumented, so durationMinutes is recorded as 0 rather than an estimate.

Run disclosure

A Codex multi-agent session generated one text-only first artifact for “diagnose unexplained high cpu usage”. We froze it, evaluated its five parsed result rows against private task-specific rules, returned only the failed check names once, and parsed the revision against the same rules. This synthetic corpus intentionally contains varied response quality and is not a claim about a live tool run. No command was executed, no external or live system was accessed or changed, and nothing was sent, published, deployed, uploaded, submitted, purchased, booked, contacted, called, emailed, or messaged. No external, live, or production action occurred. Per-case elapsed time was not instrumented during the batch session.

Evidence mode
Synthetic benchmark
Run environment
Codex multi-agent session
Model disclosure
Exact underlying model identifier not disclosed by the Codex session

03 · Verbatim input

Exact first prompt

The recorded session received the following prompt without silent additions.

Complete a bounded synthetic field test for: diagnose unexplained high cpu usage. Focus: CPU Performance.
Source scenario: The experiment will provide process samples and system metrics from a test workload with one controlled source of excess activity.
Use only these five public fictional inputs; the scoring answers are intentionally withheld:
DHC-4608-I1: System CPU rises 24%→91% at 10:10 while indexer P7 rises 8%→78%; every other process remains below 6%.
DHC-4608-I2: P7 begins scanning looped junction /docs/archive→/docs at 10:10; trace shows 41,200 repeated path visits.
DHC-4608-I3: Baseline temperature is 54°C; test stop is 92°C; faulty run reaches 90°C at minute 8.
DHC-4608-I4: Approved change excludes /docs/archive junction only; indexing of /docs/current and /docs/reference must remain.
DHC-4608-I5: Acceptance is three 15-minute runs CPU below 35%, temperature below 80°C, visits under 1,000, and searches S1-S6 6/6.
Return a concrete high-cpu causal trace with exactly five result rows, assumptions visible, and no claim that a command, message, booking, transaction, teaching session, or live-system change occurred. Evidence target: The seeded workload and before-and-after utilization traces will verify the diagnosis and proposed mitigation.

04 · Baseline preserved

First result

The first response is retained before scoring or correction.

HIGH-CPU CAUSAL TRACE DHC-4608 — FIRST FROZEN ARTIFACT
Task: diagnose unexplained high cpu usage. Evaluation focus: CPU Performance. This is a fictional, text-only artifact; it does not report a live action.
DHC-4608-R1 :: RESULT=CAUSE=replace processor hardware
DHC-4608-R1-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
DHC-4608-R2 :: RESULT=TRIGGER=normal document count
DHC-4608-R2-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
DHC-4608-R3 :: RESULT=SAFETY=continue regardless of temperature
DHC-4608-R3-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
DHC-4608-R4 :: RESULT=CHANGE=disable the entire indexer
DHC-4608-R4-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
DHC-4608-R5 :: RESULT=ACCEPT=CPU falls once
DHC-4608-R5-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
Artifact boundary: exactly five scored rows were frozen for DHC-4608; any failed row remains visible because only one correction pass is allowed.

Initial score: 0/10

05 · One pass only

Exact correction prompt

Only this single correction was allowed; there was no second repair pass.

Revise DHC-4608 once. Preserve every row that passed. Reconsider only the failed semantic checks below using the original fictional inputs; the private expected answer is not supplied:
- Identify the correlated process: input was “System CPU rises 24%→91% at 10:10 while indexer P7 rises 8%→78%; every other process remains below 6%.”; first response was “CAUSE=replace processor hardware”.
- Link the workload trigger: input was “P7 begins scanning looped junction /docs/archive→/docs at 10:10; trace shows 41,200 repeated path visits.”; first response was “TRIGGER=normal document count”.
- Respect the thermal stop: input was “Baseline temperature is 54°C; test stop is 92°C; faulty run reaches 90°C at minute 8.”; first response was “SAFETY=continue regardless of temperature”.
- Apply the bounded exclusion: input was “Approved change excludes /docs/archive junction only; indexing of /docs/current and /docs/reference must remain.”; first response was “CHANGE=disable the entire indexer”.
- Repeat performance and search checks: input was “Acceptance is three 15-minute runs CPU below 35%, temperature below 80°C, visits under 1,000, and searches S1-S6 6/6.”; first response was “ACCEPT=CPU falls once”.
Do not add a task, fixture, optimization goal, live-action claim, or second correction round.

06 · Corrected output

Corrected final result

HIGH-CPU CAUSAL TRACE DHC-4608 — AFTER ONE FAILURE-ONLY CORRECTION
Task: diagnose unexplained high cpu usage. Evaluation focus: CPU Performance. This is a fictional, text-only artifact; it does not report a live action.
DHC-4608-R1 :: RESULT=CAUSE=P7 indexer; rise70 points aligns with system rise67 points
DHC-4608-R1-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
DHC-4608-R2 :: RESULT=TRIGGER=directory junction loop; repeated visits41200
DHC-4608-R2-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
DHC-4608-R3 :: RESULT=SAFETY=stop if >=92C
DHC-4608-R3-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
DHC-4608-R4 :: RESULT=CHANGE=exclude /docs/archive junction
DHC-4608-R4-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
DHC-4608-R5 :: RESULT=ACCEPT=3x15min; CPU<35%; temp<80C; visits<1000
DHC-4608-R5-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
Artifact boundary: exactly five scored rows were frozen for DHC-4608; any failed row remains visible because only one correction pass is allowed.

Final score: 6/10

07 · Five checks, two points each

Five-check record

The first and final statuses are textual as well as color coded. Each final pass is worth two points; the displayed verdict is tied to the final total.

Five checks applied to the first and corrected results
CheckFirstFinalEvidence
Identify the correlated process Fail PassPublic fixture: System CPU rises 24%→91% at 10:10 while indexer P7 rises 8%→78%; every other process remains below 6%. Semantic rule: The timestamped process delta accounts for the system-level change. FIRST returned “CAUSE=replace processor hardware”; the private static semantic key accepts “CAUSE=P7 indexer; rise70 points aligns with system rise67 points”, so it fails. FINAL returned “CAUSE=P7 indexer; rise70 points aligns with system rise67 points”, so it passes. No live result was counted.
Link the workload trigger Fail PassPublic fixture: P7 begins scanning looped junction /docs/archive→/docs at 10:10; trace shows 41,200 repeated path visits. Semantic rule: The explicit cyclic path and repeat count explain the excess work. FIRST returned “TRIGGER=normal document count”; the private static semantic key accepts “TRIGGER=directory junction loop; repeated visits41200”, so it fails. FINAL returned “TRIGGER=directory junction loop; repeated visits41200”, so it passes. No live result was counted.
Respect the thermal stop Fail PassPublic fixture: Baseline temperature is 54°C; test stop is 92°C; faulty run reaches 90°C at minute 8. Semantic rule: The diagnostic must retain the declared temperature ceiling while interpreting the observed run. FIRST returned “SAFETY=continue regardless of temperature”; the private static semantic key accepts “SAFETY=stop if >=92C; observed90C below stop with monitoring” or “SAFETY=stop if >=92C”, so it fails. FINAL returned “SAFETY=stop if >=92C”, so it passes. No live result was counted.
Apply the bounded exclusion Fail FailPublic fixture: Approved change excludes /docs/archive junction only; indexing of /docs/current and /docs/reference must remain. Semantic rule: The causal fix removes the cyclic target without losing required search scope. FIRST returned “CHANGE=disable the entire indexer”; the private static semantic key accepts “CHANGE=exclude /docs/archive junction; retain current+reference indexing”, so it fails. FINAL returned “CHANGE=exclude /docs/archive junction”, so it fails. No live result was counted.
Repeat performance and search checks Fail FailPublic fixture: Acceptance is three 15-minute runs CPU below 35%, temperature below 80°C, visits under 1,000, and searches S1-S6 6/6. Semantic rule: Repeatability, resource limits, loop removal, and retained search behavior are joint gates. FIRST returned “ACCEPT=CPU falls once”; the private static semantic key accepts “ACCEPT=3x15min; CPU<35%; temp<80C; visits<1000; searches6/6”, so it fails. FINAL returned “ACCEPT=3x15min; CPU<35%; temp<80C; visits<1000”, so it fails. No live result was counted.
Initial0/10
Final6/10
Verdictmixed
RecommendedNo

08 · No cleanup by omission

What worked—and what failed

What worked

  • DHC-4608 preserved the exact public prompt, first artifact, failure-only correction, final artifact, and independently derived semantic check results.
  • Identify the correlated process passed because the parsed final answer matched the private fixture rule rather than merely repeating an input identifier.
  • Link the workload trigger also passed its task-specific rule with the final answer left visible.

What failed or remained weak

  • Apply the bounded exclusion still failed after the only permitted correction; its final value and expected semantic rule remain quoted in the evidence.
  • Repeat performance and search checks still failed after the only permitted correction; its final value and expected semantic rule remain quoted in the evidence.

09 · Inspectable record

Evidence notes

The seeded workload and before-and-after utilization traces will verify the diagnosis and proposed mitigation.

  • DHC-4608 stores the public five-input fixture separately from the private semantic answer strings quoted only after evaluation.
  • DHC-4608's first and final scores were recomputed from parsed RESULT rows: 0 and 3 passes multiplied by two.
  • DHC-4608 preserves every unresolved final mismatch; the source evidence plan remains unexecuted because this is a static synthetic benchmark: The seeded workload and before-and-after utilization traces will verify the diagnosis and proposed mitigation.
Download this case record

10 · Boundary of the claim

Limitations

  • DHC-4608 is a static synthetic response benchmark, not evidence that the task succeeded with a real person, organization, device, account, service, or environment.
  • DHC-4608 uses one Codex multi-agent transcript and a private deterministic fixture key; another prompt, model, evaluator, or real-world input could produce a different result.

Publication record

Published
Assigned archive date
Evidence mode
Synthetic benchmark