Computers · Evidence record
How Much Startup Delay Might AI Remove: One Verified Gap Remained
This completed synthetic Startup Performance field test asked the session to reduce a computer's startup delay, preserved an actual five-row startup delay prioritization ledger, and derived 2/10 then 8/10 from task-specific semantic checks after one failure-only correction.
- Exact prompts and outputs
- One correction only
- Synthetic inputs disclosed
01 · The assignment
The task
reduce a computer's startup delay
02 · Scope before score
Test disclosures
Input disclosure
All inputs in RSD-7216 are fictional and appear verbatim in the exact prompt. Hidden scoring answers were not shown to the response generator. No personal, production, customer, learner, or device data was used. Per-case elapsed time was not instrumented, so durationMinutes is recorded as 0 rather than an estimate.
Run disclosure
A Codex multi-agent session generated one text-only first artifact for “reduce a computer's startup delay”. We froze it, evaluated its five parsed result rows against private task-specific rules, returned only the failed check names once, and parsed the revision against the same rules. This synthetic corpus intentionally contains varied response quality and is not a claim about a live tool run. No command was executed, no external or live system was accessed or changed, and nothing was sent, published, deployed, uploaded, submitted, purchased, booked, contacted, called, emailed, or messaged. No external, live, or production action occurred. Per-case elapsed time was not instrumented during the batch session.
- Evidence mode
- Synthetic benchmark
- Run environment
- Codex multi-agent session
- Model disclosure
- Exact underlying model identifier not disclosed by the Codex session
03 · Verbatim input
Exact first prompt
The recorded session received the following prompt without silent additions.
Complete a bounded synthetic field test for: reduce a computer's startup delay. Focus: Startup Performance.
Source scenario: The experiment will ask AI to prioritize reversible startup changes on a test system with seeded background programs.
Use only these five public fictional inputs; the scoring answers are intentionally withheld:
RSD-7216-I1: Five cold boots are 86, 84, 85, 87, and 83 seconds; median is 85 seconds.
RSD-7216-I2: Items: SyncTool 19s required, OldUpdater 24s obsolete, AudioPanel 3s required, PhotoAgent 17s optional.
RSD-7216-I3: OldUpdater starts from login item L7 and scheduled task T7; both point to missing application /Old/Updater.
RSD-7216-I4: Trial policy changes OldUpdater only and stores baseline export START-BASE hash 77cb10f1.
RSD-7216-I5: Acceptance is five cold boots median below 65s, SyncTool sync pass, AudioPanel controls pass, PhotoAgent unchanged, and rollback restores L7+T7.
Return a concrete startup delay prioritization ledger with exactly five result rows, assumptions visible, and no claim that a command, message, booking, transaction, teaching session, or live-system change occurred. Evidence target: Multiple timed boots and a startup-item inventory will verify speed changes and retained functionality.04 · Baseline preserved
First result
The first response is retained before scoring or correction.
STARTUP DELAY PRIORITIZATION LEDGER RSD-7216 — FIRST FROZEN ARTIFACT
Task: reduce a computer's startup delay. Evaluation focus: Startup Performance. This is a fictional, text-only artifact; it does not report a live action.
RSD-7216-R1 :: RESULT=BASELINE=quote fastest run83s as typical
RSD-7216-R1-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
RSD-7216-R2 :: RESULT=PRIORITY=disable SyncTool because it starts first
RSD-7216-R2-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
RSD-7216-R3 :: RESULT=OLDUPDATER=remove every scheduled task
RSD-7216-R3-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
RSD-7216-R4 :: RESULT=TRIAL=freeze START-BASE 77cb10f1; disable L7+T7 only
RSD-7216-R4-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
RSD-7216-R5 :: RESULT=ACCEPT=one boot under65s
RSD-7216-R5-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
Artifact boundary: exactly five scored rows were frozen for RSD-7216; any failed row remains visible because only one correction pass is allowed.Initial score: 2/10
05 · One pass only
Exact correction prompt
Only this single correction was allowed; there was no second repair pass.
Revise RSD-7216 once. Preserve every row that passed. Reconsider only the failed semantic checks below using the original fictional inputs; the private expected answer is not supplied:
- Measure the stable baseline: input was “Five cold boots are 86, 84, 85, 87, and 83 seconds; median is 85 seconds.”; first response was “BASELINE=quote fastest run83s as typical”.
- Rank the seeded startup item: input was “Items: SyncTool 19s required, OldUpdater 24s obsolete, AudioPanel 3s required, PhotoAgent 17s optional.”; first response was “PRIORITY=disable SyncTool because it starts first”.
- Resolve the duplicate launch path: input was “OldUpdater starts from login item L7 and scheduled task T7; both point to missing application /Old/Updater.”; first response was “OLDUPDATER=remove every scheduled task”.
- Verify speed and retained function: input was “Acceptance is five cold boots median below 65s, SyncTool sync pass, AudioPanel controls pass, PhotoAgent unchanged, and rollback restores L7+T7.”; first response was “ACCEPT=one boot under65s”.
Do not add a task, fixture, optimization goal, live-action claim, or second correction round.06 · Corrected output
Corrected final result
STARTUP DELAY PRIORITIZATION LEDGER RSD-7216 — AFTER ONE FAILURE-ONLY CORRECTION
Task: reduce a computer's startup delay. Evaluation focus: Startup Performance. This is a fictional, text-only artifact; it does not report a live action.
RSD-7216-R1 :: RESULT=BASELINE=runs5; median85s; range83-87s
RSD-7216-R1-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
RSD-7216-R2 :: RESULT=PRIORITY=OldUpdater24s first; PhotoAgent17s second; retain required items
RSD-7216-R2-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
RSD-7216-R3 :: RESULT=OLDUPDATER=disable L7+T7
RSD-7216-R3-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
RSD-7216-R4 :: RESULT=TRIAL=freeze START-BASE 77cb10f1; disable L7+T7 only
RSD-7216-R4-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
RSD-7216-R5 :: RESULT=ACCEPT=5 boots median<65s; required2/2; PhotoAgent unchanged
RSD-7216-R5-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.
Artifact boundary: exactly five scored rows were frozen for RSD-7216; any failed row remains visible because only one correction pass is allowed.Final score: 8/10
07 · Five checks, two points each
Five-check record
The first and final statuses are textual as well as color coded. Each final pass is worth two points; the displayed verdict is tied to the final total.
| Check | First | Final | Evidence |
|---|---|---|---|
| Measure the stable baseline | Fail | Pass | Public fixture: Five cold boots are 86, 84, 85, 87, and 83 seconds; median is 85 seconds. Semantic rule: The median of the fixed five-run series is the comparison baseline. FIRST returned “BASELINE=quote fastest run83s as typical”; the private static semantic key accepts “BASELINE=runs5; median85s; range83-87s”, so it fails. FINAL returned “BASELINE=runs5; median85s; range83-87s”, so it passes. No live result was counted. |
| Rank the seeded startup item | Fail | Pass | Public fixture: Items: SyncTool 19s required, OldUpdater 24s obsolete, AudioPanel 3s required, PhotoAgent 17s optional. Semantic rule: The largest obsolete cost is the safest first reversible target. FIRST returned “PRIORITY=disable SyncTool because it starts first”; the private static semantic key accepts “PRIORITY=OldUpdater24s first; PhotoAgent17s second; retain required items”, so it fails. FINAL returned “PRIORITY=OldUpdater24s first; PhotoAgent17s second; retain required items”, so it passes. No live result was counted. |
| Resolve the duplicate launch path | Fail | Pass | Public fixture: OldUpdater starts from login item L7 and scheduled task T7; both point to missing application /Old/Updater. Semantic rule: Both stale launch mechanisms must be addressed without touching unrelated tasks. FIRST returned “OLDUPDATER=remove every scheduled task”; the private static semantic key accepts “OLDUPDATER=disable L7+T7; preserve inventory evidence” or “OLDUPDATER=disable L7+T7”, so it fails. FINAL returned “OLDUPDATER=disable L7+T7”, so it passes. No live result was counted. |
| Use one reversible trial | Pass | Pass | Public fixture: Trial policy changes OldUpdater only and stores baseline export START-BASE hash 77cb10f1. Semantic rule: A one-item trial preserves attribution and exact rollback. FIRST returned “TRIAL=freeze START-BASE 77cb10f1; disable L7+T7 only”; the private static semantic key accepts “TRIAL=freeze START-BASE 77cb10f1; disable L7+T7 only”, so it passes. FINAL returned “TRIAL=freeze START-BASE 77cb10f1; disable L7+T7 only”, so it passes. No live result was counted. |
| Verify speed and retained function | Fail | Fail | Public fixture: Acceptance is five cold boots median below 65s, SyncTool sync pass, AudioPanel controls pass, PhotoAgent unchanged, and rollback restores L7+T7. Semantic rule: Repeated timing, required functions, unchanged optional state, and reversal all matter. FIRST returned “ACCEPT=one boot under65s”; the private static semantic key accepts “ACCEPT=5 boots median<65s; required2/2; PhotoAgent unchanged; rollback L7+T7”, so it fails. FINAL returned “ACCEPT=5 boots median<65s; required2/2; PhotoAgent unchanged”, so it fails. No live result was counted. |
08 · No cleanup by omission
What worked—and what failed
What worked
- RSD-7216 preserved the exact public prompt, first artifact, failure-only correction, final artifact, and independently derived semantic check results.
- Measure the stable baseline passed because the parsed final answer matched the private fixture rule rather than merely repeating an input identifier.
- Rank the seeded startup item also passed its task-specific rule with the final answer left visible.
What failed or remained weak
- Verify speed and retained function still failed after the only permitted correction; its final value and expected semantic rule remain quoted in the evidence.
09 · Inspectable record
Evidence notes
Multiple timed boots and a startup-item inventory will verify speed changes and retained functionality.
- RSD-7216 stores the public five-input fixture separately from the private semantic answer strings quoted only after evaluation.
- RSD-7216's first and final scores were recomputed from parsed RESULT rows: 1 and 4 passes multiplied by two.
- RSD-7216 preserves every unresolved final mismatch; the source evidence plan remains unexecuted because this is a static synthetic benchmark: Multiple timed boots and a startup-item inventory will verify speed changes and retained functionality.
10 · Boundary of the claim
Limitations
- RSD-7216 is a static synthetic response benchmark, not evidence that the task succeeded with a real person, organization, device, account, service, or environment.
- RSD-7216 uses one Codex multi-agent transcript and a private deterministic fixture key; another prompt, model, evaluator, or real-world input could produce a different result.