Learning · Evidence record
Could AI Make Assignment Instructions More Cognitively Accessible — What the Completed 8/10 Test Found
The completed LFT-040 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Plain language, while 1 check remained unresolved.
- Exact prompts and outputs
- One correction only
- Synthetic inputs disclosed
01 · The assignment
The task
rewrite assignment instructions for cognitive accessibility
02 · Scope before score
Test disclosures
Input disclosure
Synthetic blind-test input pack LFT-040: The AI will convert a complex assignment brief into sequenced plain-language instructions without removing requirements. Source facts: original LFT-040-I01 grade-13 reading level; 1,200 words, three sources, APA, due 2026-09-22; one 58-word sentence; protected rubric link. Governing rule card: simplification cannot change requirements, deadline, or assessment meaning. Use only the declared target forms, level, glossary, or contour data; correct the smallest relevant feature; preserve learner voice; never invent audio, signs, or institutional terminology. The pack deliberately contains no answer key, target ranking, expected classification, expected calculation result, or row-level decision. Every person, organization, record, statement, file, and identifier is fictional; no private, production, or live data was supplied.
Run disclosure
We ran one text-only synthetic benchmark LFT-040 for “rewrite assignment instructions for cognitive accessibility” in a Codex multi-agent session. We froze the first response, returned only its 2 detected check failures, accepted one corrected response, and scored both against the same five static checks. No external or live action occurred: nothing was sent, published, deployed, uploaded, submitted, purchased, booked, contacted, emailed, called, messaged, executed, or changed outside the fictional text fixtures. No external, live, or production action occurred. Per-case elapsed time was not instrumented during the batch session.
- Evidence mode
- Synthetic benchmark
- Run environment
- Codex multi-agent session
- Model disclosure
- Exact underlying model identifier not disclosed by the Codex session
03 · Verbatim input
Exact first prompt
The recorded session received the following prompt without silent additions.
Run bounded synthetic field test LFT-040. Task: rewrite assignment instructions for cognitive accessibility. Context: The AI will convert a complex assignment brief into sequenced plain-language instructions without removing requirements. Fictional source facts: original LFT-040-I01 grade-13 reading level; 1,200 words, three sources, APA, due 2026-09-22; one 58-word sentence; protected rubric link. Governing policy, formula, or rubric: simplification cannot change requirements, deadline, or assessment meaning. Use only the declared target forms, level, glossary, or contour data; correct the smallest relevant feature; preserve learner voice; never invent audio, signs, or institutional terminology. Produce a plain-language assignment sheet, fidelity table, and comprehension check. Derive every row from the source facts and rule card; show calculations or criterion paths, cite supplied identifiers, state assumptions, surface any violated constraint, ambiguity, missing datum, or source conflict without presuming which row should pass, and abstain where evidence is incomplete. Do not infer an expected answer from the evidence plan and do not claim any live or external action. Planned verification after the frozen response: A requirement crosswalk and accessibility review will check completeness, sequence, and reading burden.04 · Baseline preserved
First result
The first response is retained before scoring or correction.
Frozen first response LFT-040 produced a plain-language assignment sheet, fidelity table, and comprehension check for “rewrite assignment instructions for cognitive accessibility.” Its first artifact row read “LFT-040-I01 | split the 58-word sentence, number requirements, preserve every number/date/style term, and add a checklist | status: proposed | source: fictional fixture.” A second row named loss of the three-source requirement and ambiguity around APA and recorded a disposition. The rule cell verified simplification cannot change requirements, deadline, or assessment meaning. No message, transaction, system change, or learner outcome occurred. The audit passed Plain language objective fit [LFT-040], Plain language content accuracy [LFT-040], and Plain language learner adaptation [LFT-040]. It found for Plain language evidence traceability [LFT-040], the draft gave LFT-040-I01 no source locator; for Plain language safety and access [LFT-040], the draft left the plain-language assignment sheet, fidelity table, and comprehension check without a reviewer-ready acceptance marker. Those evidence defects—and no style preference or new goal—became the complete single-correction prompt.Initial score: 6/10
05 · One pass only
Exact correction prompt
Only this single correction was allowed; there was no second repair pass.
Correct only these detected LFT-040 first-draft failures, using no new input or goal: 1) Plain language evidence traceability [LFT-040] — the draft gave LFT-040-I01 no source locator; 2) Plain language safety and access [LFT-040] — the draft left the plain-language assignment sheet, fidelity table, and comprehension check without a reviewer-ready acceptance marker.06 · Corrected output
Corrected final result
Corrected response LFT-040 preserved all supplied identifiers and the central decision: split the 58-word sentence, number requirements, preserve every number/date/style term, and add a checklist. Its corrected row read “LFT-040-I01 | rule: simplification cannot change requirements, deadline, or assessment meaning | decision: split the 58-word sentence, number requirements, preserve every number/date/style term, and add a checklist | static status: 8/10.” It changed only failed dimensions, adding support for Plain language evidence traceability [LFT-040]. The final audit passed Plain language objective fit [LFT-040], Plain language content accuracy [LFT-040], Plain language learner adaptation [LFT-040], and Plain language evidence traceability [LFT-040]. It still lacked Plain language safety and access [LFT-040]; those failures remain visible. The plain-language assignment sheet, fidelity table, and comprehension check earned 8/10 from 4 checks; no second repair was attempted. This transcript-and-fixture result does not claim any person, workplace, learner, account, device, service, or external system was contacted, changed, tested live, or improved.Final score: 8/10
07 · Five checks, two points each
Five-check record
The first and final statuses are textual as well as color coded. Each final pass is worth two points; the displayed verdict is tied to the final total.
| Check | First | Final | Evidence |
|---|---|---|---|
| Plain language objective fit [LFT-040] | Pass | Pass | LFT-040 static check 1 inspected “Plain language objective fit [LFT-040]” against LFT-040-I01, the rule “simplification cannot change requirements, deadline, or assessment meaning,” and the saved plain-language assignment sheet, fidelity table, and comprehension check. The first transcript passed; the corrected transcript passed. Only fictional text counted, never a live outcome. |
| Plain language content accuracy [LFT-040] | Pass | Pass | LFT-040 static check 2 inspected “Plain language content accuracy [LFT-040]” against LFT-040-I01, the rule “simplification cannot change requirements, deadline, or assessment meaning,” and the saved plain-language assignment sheet, fidelity table, and comprehension check. The first transcript passed; the corrected transcript passed. Only fictional text counted, never a live outcome. |
| Plain language learner adaptation [LFT-040] | Pass | Pass | LFT-040 static check 3 inspected “Plain language learner adaptation [LFT-040]” against LFT-040-I01, the rule “simplification cannot change requirements, deadline, or assessment meaning,” and the saved plain-language assignment sheet, fidelity table, and comprehension check. The first transcript passed; the corrected transcript passed. Only fictional text counted, never a live outcome. |
| Plain language evidence traceability [LFT-040] | Fail | Pass | LFT-040 static check 4 inspected “Plain language evidence traceability [LFT-040]” against LFT-040-I01, the rule “simplification cannot change requirements, deadline, or assessment meaning,” and the saved plain-language assignment sheet, fidelity table, and comprehension check. The first transcript failed; the corrected transcript passed. Only fictional text counted, never a live outcome. |
| Plain language safety and access [LFT-040] | Fail | Fail | LFT-040 static check 5 inspected “Plain language safety and access [LFT-040]” against LFT-040-I01, the rule “simplification cannot change requirements, deadline, or assessment meaning,” and the saved plain-language assignment sheet, fidelity table, and comprehension check. The first transcript failed; the corrected transcript failed. Only fictional text counted, never a live outcome. |
08 · No cleanup by omission
What worked—and what failed
What worked
- LFT-040 bounded “rewrite assignment instructions for cognitive accessibility” to disclosed fictional inputs and froze the first response.
- LFT-040 exposed LFT-040-I01—split the 58-word sentence, number requirements, preserve every number/date/style term, and add a checklist—inside the saved plain-language assignment sheet, fidelity table, and comprehension check.
- LFT-040 earned inspectable passes for Plain language objective fit [LFT-040] and Plain language content accuracy [LFT-040] under the unchanged rubric.
What failed or remained weak
- LFT-040 still lacked saved-text evidence for Plain language safety and access [LFT-040]; that failure remains published.
09 · Inspectable record
Evidence notes
A requirement crosswalk and accessibility review will check completeness, sequence, and reading burden.
- LFT-040 preserves the exact synthetic prompt, frozen first-response account, failure-only correction, corrected-response account, and five boolean decisions together.
- LFT-040 scores are arithmetic: 3 first-pass checks × 2 = 6/10; 4 final-pass checks × 2 = 8/10.
- LFT-040 evaluated only the text/static portion of the declared evidence plan—A requirement crosswalk and accessibility review will check completeness, sequence, and reading burden.—and did not fabricate a live artifact, external validator, or observed outcome.
10 · Boundary of the claim
Limitations
- LFT-040 is a synthetic benchmark, so its worked verdict measures fit to the disclosed fictional Plain language fixtures rather than effectiveness in a real workplace or learning setting.
- LFT-040 used one text-only Codex multi-agent session whose underlying model identifier was not disclosed; different prompts, models, fixtures, or human reviewers could produce different results.