{
  "category": "learning",
  "slug": "learning-adult-budget-literacy",
  "title": "Would an AI Budget Coach Fit an Adult Learner's Scenario: A Failed Synthetic Benchmark at 4/10",
  "task": "teach budgeting through an adult learner's scenario",
  "excerpt": "The completed LFT-011 synthetic field test finished at 4/10 and was not recommended: only two of five Financial literacy checks passed after the permitted correction.",
  "tool": "Codex multi-agent session",
  "model": "Exact underlying model identifier not disclosed by the Codex session",
  "publishedAt": "2026-03-08T18:00:00+08:00",
  "durationMinutes": 0,
  "testMode": "Synthetic benchmark",
  "inputDisclosure": "Synthetic blind-test input pack LFT-011: An adult learner will build a monthly budget from a fictional household scenario with changing expenses. Source facts: income $3,850; rent $1,420; utilities $210; food $520; transport $330; debt $480; savings $400; repair adds $650 month two. Governing rule card: income minus categorized expenses with learner-chosen trade-offs. Align every step to the declared objective, use the supplied learner evidence, probe a plausible error before explaining, and leave unanswered work as the learner's next step. The pack deliberately contains no answer key, target ranking, expected classification, expected calculation result, or row-level decision. Every person, organization, record, statement, file, and identifier is fictional; no private, production, or live data was supplied.",
  "runDisclosure": "We ran one text-only synthetic benchmark LFT-011 for “teach budgeting through an adult learner's scenario” in a Codex multi-agent session. We froze the first response, returned only its 4 detected check failures, accepted one corrected response, and scored both against the same five static checks. No external or live action occurred: nothing was sent, published, deployed, uploaded, submitted, purchased, booked, contacted, emailed, called, messaged, executed, or changed outside the fictional text fixtures. No external, live, or production action occurred. Per-case elapsed time was not instrumented during the batch session.",
  "prompt": "Run bounded synthetic field test LFT-011. Task: teach budgeting through an adult learner's scenario. Context: An adult learner will build a monthly budget from a fictional household scenario with changing expenses. Fictional source facts: income $3,850; rent $1,420; utilities $210; food $520; transport $330; debt $480; savings $400; repair adds $650 month two. Governing policy, formula, or rubric: income minus categorized expenses with learner-chosen trade-offs. Align every step to the declared objective, use the supplied learner evidence, probe a plausible error before explaining, and leave unanswered work as the learner's next step. Produce a monthly household budget, change log, and trade-off explanation. Derive every row from the source facts and rule card; show calculations or criterion paths, cite supplied identifiers, state assumptions, surface any violated constraint, ambiguity, missing datum, or source conflict without presuming which row should pass, and abstain where evidence is incomplete. Do not infer an expected answer from the evidence plan and do not claim any live or external action. Planned verification after the frozen response: The saved budget and decision log will verify arithmetic accuracy and responses to each changed constraint.",
  "firstResult": "Frozen first response LFT-011 produced a monthly household budget, change log, and trade-off explanation for “teach budgeting through an adult learner's scenario.” Its first artifact row read “LFT-011-B02 | calculate month-one remainder $490, absorb repair through explicit discretionary/savings changes, and protect debt $480 | status: proposed | source: fictional fixture.” A second row named the month-two $160 shortfall and protected debt minimum and recorded a disposition. The rule cell mentioned without verifying income minus categorized expenses with learner-chosen trade-offs. No message, transaction, system change, or learner outcome occurred. The audit passed Financial literacy learner adaptation [LFT-011]. It found for Financial literacy objective fit [LFT-011], the draft did not link LFT-011-B02 to the full task boundary; for Financial literacy content accuracy [LFT-011], the draft mentioned but did not verify income minus categorized expenses with learner-chosen trade-offs; for Financial literacy evidence traceability [LFT-011], the draft gave LFT-011-B02 no source locator; for Financial literacy safety and access [LFT-011], the draft left the monthly household budget, change log, and trade-off explanation without a reviewer-ready acceptance marker. Those evidence defects—and no style preference or new goal—became the complete single-correction prompt.",
  "correctionPrompt": "Correct only these detected LFT-011 first-draft failures, using no new input or goal: 1) Financial literacy objective fit [LFT-011] — the draft did not link LFT-011-B02 to the full task boundary; 2) Financial literacy content accuracy [LFT-011] — the draft mentioned but did not verify income minus categorized expenses with learner-chosen trade-offs; 3) Financial literacy evidence traceability [LFT-011] — the draft gave LFT-011-B02 no source locator; 4) Financial literacy safety and access [LFT-011] — the draft left the monthly household budget, change log, and trade-off explanation without a reviewer-ready acceptance marker.",
  "finalResult": "Corrected response LFT-011 preserved all supplied identifiers and the central decision: calculate month-one remainder $490, absorb repair through explicit discretionary/savings changes, and protect debt $480. Its corrected row read “LFT-011-B02 | rule: income minus categorized expenses with learner-chosen trade-offs | decision: calculate month-one remainder $490, absorb repair through explicit discretionary/savings changes, and protect debt $480 | static status: 4/10.” It changed only failed dimensions, adding support for Financial literacy evidence traceability [LFT-011]. The final audit passed Financial literacy learner adaptation [LFT-011] and Financial literacy evidence traceability [LFT-011]. It still lacked Financial literacy objective fit [LFT-011], Financial literacy content accuracy [LFT-011], and Financial literacy safety and access [LFT-011]; those failures remain visible. The monthly household budget, change log, and trade-off explanation earned 4/10 from 2 checks; no second repair was attempted. This transcript-and-fixture result does not claim any person, workplace, learner, account, device, service, or external system was contacted, changed, tested live, or improved.",
  "checks": [
    {
      "name": "Financial literacy objective fit [LFT-011]",
      "firstPass": false,
      "finalPass": false,
      "evidence": "LFT-011 static check 1 inspected “Financial literacy objective fit [LFT-011]” against LFT-011-B02, the rule “income minus categorized expenses with learner-chosen trade-offs,” and the saved monthly household budget, change log, and trade-off explanation. The first transcript failed; the corrected transcript failed. Only fictional text counted, never a live outcome."
    },
    {
      "name": "Financial literacy content accuracy [LFT-011]",
      "firstPass": false,
      "finalPass": false,
      "evidence": "LFT-011 static check 2 inspected “Financial literacy content accuracy [LFT-011]” against LFT-011-B02, the rule “income minus categorized expenses with learner-chosen trade-offs,” and the saved monthly household budget, change log, and trade-off explanation. The first transcript failed; the corrected transcript failed. Only fictional text counted, never a live outcome."
    },
    {
      "name": "Financial literacy learner adaptation [LFT-011]",
      "firstPass": true,
      "finalPass": true,
      "evidence": "LFT-011 static check 3 inspected “Financial literacy learner adaptation [LFT-011]” against LFT-011-B02, the rule “income minus categorized expenses with learner-chosen trade-offs,” and the saved monthly household budget, change log, and trade-off explanation. The first transcript passed; the corrected transcript passed. Only fictional text counted, never a live outcome."
    },
    {
      "name": "Financial literacy evidence traceability [LFT-011]",
      "firstPass": false,
      "finalPass": true,
      "evidence": "LFT-011 static check 4 inspected “Financial literacy evidence traceability [LFT-011]” against LFT-011-B02, the rule “income minus categorized expenses with learner-chosen trade-offs,” and the saved monthly household budget, change log, and trade-off explanation. The first transcript failed; the corrected transcript passed. Only fictional text counted, never a live outcome."
    },
    {
      "name": "Financial literacy safety and access [LFT-011]",
      "firstPass": false,
      "finalPass": false,
      "evidence": "LFT-011 static check 5 inspected “Financial literacy safety and access [LFT-011]” against LFT-011-B02, the rule “income minus categorized expenses with learner-chosen trade-offs,” and the saved monthly household budget, change log, and trade-off explanation. The first transcript failed; the corrected transcript failed. Only fictional text counted, never a live outcome."
    }
  ],
  "initialScore": 2,
  "score": 4,
  "verdict": "failed",
  "recommended": false,
  "whatWorked": [
    "LFT-011 bounded “teach budgeting through an adult learner's scenario” to disclosed fictional inputs and froze the first response.",
    "LFT-011 exposed LFT-011-B02—calculate month-one remainder $490, absorb repair through explicit discretionary/savings changes, and protect debt $480—inside the saved monthly household budget, change log, and trade-off explanation."
  ],
  "whatFailed": [
    "LFT-011 still lacked saved-text evidence for Financial literacy objective fit [LFT-011]; that failure remains published.",
    "LFT-011 still lacked saved-text evidence for Financial literacy content accuracy [LFT-011]; that failure remains published.",
    "LFT-011 still lacked saved-text evidence for Financial literacy safety and access [LFT-011]; that failure remains published."
  ],
  "evidencePlan": "The saved budget and decision log will verify arithmetic accuracy and responses to each changed constraint.",
  "evidenceNotes": [
    "LFT-011 preserves the exact synthetic prompt, frozen first-response account, failure-only correction, corrected-response account, and five boolean decisions together.",
    "LFT-011 scores are arithmetic: 1 first-pass checks × 2 = 2/10; 2 final-pass checks × 2 = 4/10.",
    "LFT-011 evaluated only the text/static portion of the declared evidence plan—The saved budget and decision log will verify arithmetic accuracy and responses to each changed constraint.—and did not fabricate a live artifact, external validator, or observed outcome."
  ],
  "limitations": [
    "LFT-011 is a synthetic benchmark, so its failed verdict measures fit to the disclosed fictional Financial literacy fixtures rather than effectiveness in a real workplace or learning setting.",
    "LFT-011 used one text-only Codex multi-agent session whose underlying model identifier was not disclosed; different prompts, models, fixtures, or human reviewers could produce different results."
  ]
}
