Evidence guide · 41 published records

AI for teaching and study support

Compare tests of explanations, practice plans, feedback, misconceptions, and learning sequences across bounded lessons.

The decision this guide supports

Can AI support the reasoning process without skipping the misconception, giving away the answer, or overstating certainty?

Published records
41
Average first score
5.2/10
Average final score
7.6/10
Average gain
+2.4
Worked / mixed / failed
29 / 8 / 4

Averages use 41 records with a disclosed first score. The set contains 41 synthetic benchmarks and 0 file-backed tests; it is a task collection, not a representative model leaderboard.

What the records compare

Evidence before a recommendation.

Concept accuracy, learner steps, feedback quality, answer leakage, and checks that match the stated learning goal.

Across this released set, 4 records failed and 8 remained mixed after the single correction. Those outcomes stay in the guide because a useful decision needs the misses as well as the wins.

  1. Name the learner level and the exact misconception or goal.
  2. Check examples and explanations against a trusted reference.
  3. Keep assessment and safeguarding decisions with a qualified person.

Representative evidence

Open the prompts and checks.

The sample deliberately includes different verdicts when available. Every card opens to the full first result, correction, final result, and evidence boundary.