Evidence guide · 41 published records
AI for teaching and study support
Compare tests of explanations, practice plans, feedback, misconceptions, and learning sequences across bounded lessons.
The decision this guide supports
Can AI support the reasoning process without skipping the misconception, giving away the answer, or overstating certainty?
- Published records
- 41
- Average first score
- 5.2/10
- Average final score
- 7.6/10
- Average gain
- +2.4
- Worked / mixed / failed
- 29 / 8 / 4
Averages use 41 records with a disclosed first score. The set contains 41 synthetic benchmarks and 0 file-backed tests; it is a task collection, not a representative model leaderboard.
What the records compare
Evidence before a recommendation.
Concept accuracy, learner steps, feedback quality, answer leakage, and checks that match the stated learning goal.
Across this released set, 4 records failed and 8 remained mixed after the single correction. Those outcomes stay in the guide because a useful decision needs the misses as well as the wins.
- Name the learner level and the exact misconception or goal.
- Check examples and explanations against a trusted reference.
- Keep assessment and safeguarding decisions with a qualified person.
Representative evidence
Open the prompts and checks.
The sample deliberately includes different verdicts when available. Every card opens to the full first result, correction, final result, and evidence boundary.
Completed field testSynthetic benchmark
Is an AI-Adapted Diagram Lesson Accessible Without Sight
The completed LFT-028 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Nonvisual access, while 0 checks remained unresolved.
Completed field testSynthetic benchmark
Using Branching Scenarios for AI-Guided Workplace-Safety Training
The completed LFT-042 synthetic field test stopped at 6/10: three of five Safety training checks passed after one correction, but Safety training learner adaptation [LFT-042] and Safety training evidence traceability [LFT-042] remained unsupported.
Completed field testSynthetic benchmark
Teaching the Product Rule Through AI-Led Error Analysis
The completed LFT-035 synthetic field test finished at 4/10 and was not recommended: only two of five Calculus errors checks passed after the permitted correction.
Completed field testSynthetic benchmark
Planning an Ecosystem Inquiry Lesson with AI
The completed LFT-003 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Lesson planning, while 0 checks remained unresolved.
Completed field testSynthetic benchmark
Turning Transit Maps into Graduated Reading Practice
The completed LFT-068 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Map Literacy, while 0 checks remained unresolved.
Completed field testSynthetic benchmark
Map Route Distances with AI to Learn Scale
The completed LFT-033 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Map reasoning, while 0 checks remained unresolved.