{
  "category": "computers",
  "slug": "computers-extend-laptop-battery",
  "title": "Battery Life Without a Broken Workflow: An AI Tuning Brief: Three Semantic Checks Still Failed",
  "task": "extend laptop battery life without crippling usability",
  "excerpt": "This completed synthetic Battery Life field test asked the session to extend laptop battery life without crippling usability, preserved an actual five-row laptop battery usability trial, and derived 0/10 then 4/10 from task-specific semantic checks after one failure-only correction.",
  "tool": "Codex multi-agent session",
  "model": "Exact underlying model identifier not disclosed by the Codex session",
  "publishedAt": "2026-04-10T14:00:00+08:00",
  "durationMinutes": 0,
  "testMode": "Synthetic benchmark",
  "inputDisclosure": "All inputs in ELB-5103 are fictional and appear verbatim in the exact prompt. Hidden scoring answers were not shown to the response generator. No personal, production, customer, learner, or device data was used. Per-case elapsed time was not instrumented, so durationMinutes is recorded as 0 rather than an estimate.",
  "runDisclosure": "A Codex multi-agent session generated one text-only first artifact for “extend laptop battery life without crippling usability”. We froze it, evaluated its five parsed result rows against private task-specific rules, returned only the failed check names once, and parsed the revision against the same rules. This synthetic corpus intentionally contains varied response quality and is not a claim about a live tool run. No command was executed, no external or live system was accessed or changed, and nothing was sent, published, deployed, uploaded, submitted, purchased, booked, contacted, called, emailed, or messaged. No external, live, or production action occurred. Per-case elapsed time was not instrumented during the batch session.",
  "prompt": "Complete a bounded synthetic field test for: extend laptop battery life without crippling usability. Focus: Battery Life.\nSource scenario: The experiment will ask AI to prioritize reversible power settings for a fixed mix of browsing, video, and document work.\nUse only these five public fictional inputs; the scoring answers are intentionally withheld:\nELB-5103-I1: Full-charge capacity is 61 Wh versus design 72 Wh; cycle count is 410 and no service warning is present.\nELB-5103-I2: Trial is 90 min browsing, 60 min video, and 90 min documents with brightness 60%; baseline uses 43 Wh.\nELB-5103-I3: Fixture estimates: background sync 4.8 Wh, display 3.1 Wh, keyboard light 0.6 Wh; sync may defer but must finish by 18:00.\nELB-5103-I4: Video must remain 1080p, brightness at least 45%, page response below 300 ms, and all sync work complete by 18:00.\nELB-5103-I5: Acceptance is three identical trials using at most 35 Wh, response below 300 ms, background work complete, and rollback restores baseline settings.\nReturn a concrete laptop battery usability trial with exactly five result rows, assumptions visible, and no claim that a command, message, booking, transaction, teaching session, or live-system change occurred. Evidence target: Repeated workload runs will verify runtime, responsiveness, display behavior, and background-task completion.",
  "firstResult": "LAPTOP BATTERY USABILITY TRIAL ELB-5103 — FIRST FROZEN ARTIFACT\nTask: extend laptop battery life without crippling usability. Evaluation focus: Battery Life. This is a fictional, text-only artifact; it does not report a live action.\nELB-5103-R1 :: RESULT=HEALTH=118% because 72/61\nELB-5103-R1-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nELB-5103-R2 :: RESULT=WORKLOAD=screen-off idle test\nELB-5103-R2-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nELB-5103-R3 :: RESULT=PRIORITY=disable document autosave\nELB-5103-R3-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nELB-5103-R4 :: RESULT=USABILITY=brightness20% and video480p\nELB-5103-R4-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nELB-5103-R5 :: RESULT=ACCEPT=one longer runtime estimate\nELB-5103-R5-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nArtifact boundary: exactly five scored rows were frozen for ELB-5103; any failed row remains visible because only one correction pass is allowed.",
  "correctionPrompt": "Revise ELB-5103 once. Preserve every row that passed. Reconsider only the failed semantic checks below using the original fictional inputs; the private expected answer is not supplied:\n- Calculate the battery-health baseline: input was “Full-charge capacity is 61 Wh versus design 72 Wh; cycle count is 410 and no service warning is present.”; first response was “HEALTH=118% because 72/61”.\n- Use the fixed mixed workload: input was “Trial is 90 min browsing, 60 min video, and 90 min documents with brightness 60%; baseline uses 43 Wh.”; first response was “WORKLOAD=screen-off idle test”.\n- Prioritize measured reversible savings: input was “Fixture estimates: background sync 4.8 Wh, display 3.1 Wh, keyboard light 0.6 Wh; sync may defer but must finish by 18:00.”; first response was “PRIORITY=disable document autosave”.\n- Preserve usability constraints: input was “Video must remain 1080p, brightness at least 45%, page response below 300 ms, and all sync work complete by 18:00.”; first response was “USABILITY=brightness20% and video480p”.\n- Compare repeated workload results: input was “Acceptance is three identical trials using at most 35 Wh, response below 300 ms, background work complete, and rollback restores baseline settings.”; first response was “ACCEPT=one longer runtime estimate”.\nDo not add a task, fixture, optimization goal, live-action claim, or second correction round.",
  "finalResult": "LAPTOP BATTERY USABILITY TRIAL ELB-5103 — AFTER ONE FAILURE-ONLY CORRECTION\nTask: extend laptop battery life without crippling usability. Evaluation focus: Battery Life. This is a fictional, text-only artifact; it does not report a live action.\nELB-5103-R1 :: RESULT=HEALTH=61/72=84.7%; cycles410; no service warning\nELB-5103-R1-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nELB-5103-R2 :: RESULT=WORKLOAD=240min total; brightness60%; baseline43Wh\nELB-5103-R2-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nELB-5103-R3 :: RESULT=PRIORITY=defer sync4.8Wh first; display3.1Wh second\nELB-5103-R3-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nELB-5103-R4 :: RESULT=USABILITY=video1080p; brightness>=45%; response<300ms\nELB-5103-R4-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nELB-5103-R5 :: RESULT=ACCEPT=3 trials<=35Wh; response<300ms; sync complete\nELB-5103-R5-NOTE :: The proposed technical step is static and bounded; no command output or successful device change is invented.\nArtifact boundary: exactly five scored rows were frozen for ELB-5103; any failed row remains visible because only one correction pass is allowed.",
  "checks": [
    {
      "name": "Calculate the battery-health baseline",
      "firstPass": false,
      "finalPass": true,
      "evidence": "Public fixture: Full-charge capacity is 61 Wh versus design 72 Wh; cycle count is 410 and no service warning is present. Semantic rule: Capacity ratio, cycle count, and status are distinct baseline facts. FIRST returned “HEALTH=118% because 72/61”; the private static semantic key accepts “HEALTH=61/72=84.7%; cycles410; no service warning”, so it fails. FINAL returned “HEALTH=61/72=84.7%; cycles410; no service warning”, so it passes. No live result was counted."
    },
    {
      "name": "Use the fixed mixed workload",
      "firstPass": false,
      "finalPass": true,
      "evidence": "Public fixture: Trial is 90 min browsing, 60 min video, and 90 min documents with brightness 60%; baseline uses 43 Wh. Semantic rule: Runtime claims must use the disclosed four-hour representative workload. FIRST returned “WORKLOAD=screen-off idle test”; the private static semantic key accepts “WORKLOAD=240min total; brightness60%; baseline43Wh”, so it fails. FINAL returned “WORKLOAD=240min total; brightness60%; baseline43Wh”, so it passes. No live result was counted."
    },
    {
      "name": "Prioritize measured reversible savings",
      "firstPass": false,
      "finalPass": false,
      "evidence": "Public fixture: Fixture estimates: background sync 4.8 Wh, display 3.1 Wh, keyboard light 0.6 Wh; sync may defer but must finish by 18:00. Semantic rule: The measured savings and functional deadline determine the order. FIRST returned “PRIORITY=disable document autosave”; the private static semantic key accepts “PRIORITY=defer sync4.8Wh first; display3.1Wh second; keyboard0.6Wh third”, so it fails. FINAL returned “PRIORITY=defer sync4.8Wh first; display3.1Wh second”, so it fails. No live result was counted."
    },
    {
      "name": "Preserve usability constraints",
      "firstPass": false,
      "finalPass": false,
      "evidence": "Public fixture: Video must remain 1080p, brightness at least 45%, page response below 300 ms, and all sync work complete by 18:00. Semantic rule: The task explicitly disallows battery gains that violate the four usability thresholds. FIRST returned “USABILITY=brightness20% and video480p”; the private static semantic key accepts “USABILITY=video1080p; brightness>=45%; response<300ms; sync by18:00”, so it fails. FINAL returned “USABILITY=video1080p; brightness>=45%; response<300ms”, so it fails. No live result was counted."
    },
    {
      "name": "Compare repeated workload results",
      "firstPass": false,
      "finalPass": false,
      "evidence": "Public fixture: Acceptance is three identical trials using at most 35 Wh, response below 300 ms, background work complete, and rollback restores baseline settings. Semantic rule: Repeated energy, responsiveness, task-completion, and rollback evidence all apply. FIRST returned “ACCEPT=one longer runtime estimate”; the private static semantic key accepts “ACCEPT=3 trials<=35Wh; response<300ms; sync complete; rollback exact”, so it fails. FINAL returned “ACCEPT=3 trials<=35Wh; response<300ms; sync complete”, so it fails. No live result was counted."
    }
  ],
  "initialScore": 0,
  "score": 4,
  "verdict": "failed",
  "recommended": false,
  "whatWorked": [
    "ELB-5103 preserved the exact public prompt, first artifact, failure-only correction, final artifact, and independently derived semantic check results.",
    "Calculate the battery-health baseline passed because the parsed final answer matched the private fixture rule rather than merely repeating an input identifier.",
    "Use the fixed mixed workload also passed its task-specific rule with the final answer left visible."
  ],
  "whatFailed": [
    "Prioritize measured reversible savings still failed after the only permitted correction; its final value and expected semantic rule remain quoted in the evidence.",
    "Preserve usability constraints still failed after the only permitted correction; its final value and expected semantic rule remain quoted in the evidence.",
    "Compare repeated workload results still failed after the only permitted correction; its final value and expected semantic rule remain quoted in the evidence."
  ],
  "evidencePlan": "Repeated workload runs will verify runtime, responsiveness, display behavior, and background-task completion.",
  "evidenceNotes": [
    "ELB-5103 stores the public five-input fixture separately from the private semantic answer strings quoted only after evaluation.",
    "ELB-5103's first and final scores were recomputed from parsed RESULT rows: 0 and 2 passes multiplied by two.",
    "ELB-5103 preserves every unresolved final mismatch; the source evidence plan remains unexecuted because this is a static synthetic benchmark: Repeated workload runs will verify runtime, responsiveness, display behavior, and background-task completion."
  ],
  "limitations": [
    "ELB-5103 is a static synthetic response benchmark, not evidence that the task succeeded with a real person, organization, device, account, service, or environment.",
    "ELB-5103 uses one Codex multi-agent transcript and a private deterministic fixture key; another prompt, model, evaluator, or real-world input could produce a different result."
  ]
}
