RubricMark · Study Lab

Can you use what you learned without AI help?

Unguarded chat lifted practice by about 48% and cut the later unassisted exam by about 17%. A hint-only tutor kept the practice gain and largely removed the harm. Switch the chart to a 45-day quiz that found the same direction.

A chat that writes the next line feels like progress. The score that matters is the one you get after the window is gone. Bastani and colleagues ran that comparison in high-school maths with nearly a thousand students and two different tutors built on the same model.

One tutor behaved like an ordinary chat: students could ask for the solution and copy it. The other was prompted to hint, to use teacher-written solutions and common mistakes, and not to hand over the finished line. A third group had no AI. Everyone later sat an exam with the tool removed.

Practice up, exam down — unless the tool withholds

Against the no-AI group, unguarded chat improved assisted practice by about 48%. The hint-only tutor improved practice by about 127%. On the unassisted exam, unguarded chat cut performance by about 17%. The hint-only arm was statistically indistinguishable from no AI. The point estimate was slightly negative (about four thousandths on a 0–1 scale), which is why the chart draws that bar at zero rather than inventing a gain. The paper’s own line is that guardrails largely removed the harm and did not show a positive exam effect.

Students with the open chat often asked for solutions and copied them. Students with the hinting tutor asked for help more often than for the finished answer. The practice score, in other words, measured the tool as much as the student.

[[chart:ld-ai-guardrails]]
Bastani et al. 2025 — change versus no AI
ArmAssisted practiceUnassisted exam
No AIbaselinebaseline
Unguarded chatabout +48%about −17%
Hint-only tutorabout +127%indistinguishable from no AI

A second study, a longer delay

Switch the chart to “45-day quiz.” Barcaui’s 2025 trial compared ChatGPT-assisted study with traditional study and found lower retention at 45 days: 57% correct versus 68%. Different design, same warning. Fluency while the model is open is not retention after it leaves.

None of this says “AI tutors never work.” Kestin and colleagues, in Scientific Reports (2025), found a content-locked, pedagogy-matched physics tutor beat in-class active learning on an immediate post-test at Harvard (N=194). Design is the variable. An open answer box is a different product from a tutor that withholds the line until you have tried.

These are not RubricMark trials. We do not claim a percentage lift for our own cards. We use the published pattern as a design rule: AI may draft a deck or mark after you write. It does not flip the card or sit the paper.

A rule you can apply tonight

  1. Attempt the question with the chat closed.
  2. If you are stuck, ask for a hint that names the first move, not the last line.
  3. Write your own solution before you compare.
  4. Put the miss on a card or a redrill. Do not save the chat transcript as the study record.

That is the loop on Study Lab: retrieve until the back is yours. Marks after you write live at assignment feedback. A named timed paper, when one exists, is Practice Hub — composed from the exam bank, not invented in the moment by a chat.

What to stop measuring

Stop measuring the evening by how quickly the homework page emptied. In the Bastani study the open chat emptied the practice set and the exam got worse. A useful log is three lines: what you attempted with the chat closed, what you still could not do, and what you put on a card. If the middle line is empty every night, you are choosing work that is too easy or you are peeking. Either way the unassisted column will not move.

Parents comparing two students should not compare assisted practice scores. Those scores include the tool. Compare a second attempt with the tool shut, or a due-card pile that actually cleared. RubricMark does not publish a “our students gained 17%” figure, and this article should not be read as one. The 17% is the published harm of unguarded chat in that high-school maths field experiment. The design response is attempt first, hint before answer, marks after you write.

Kestin’s physics result is the other half of the same sentence. A tutor that is locked to the course and matched to the pedagogy can help on an immediate test. That is not permission to paste the worksheet into a general chat. If you want a model, write yours first. Then compare. Then retrieve the step you missed, on a later day, with the comparison closed.

Related