A deep research pass with adversarial verification. Every claim here survived three independent agents trying to knock it down. Fifteen claims did not survive, and those are listed too.
| Question | Status |
|---|---|
| Parent dashboard and progress reporting | Answered |
| Child motivation systems, with critique | Answered |
| Dual parent/child mode patterns | Not answered |
| Typography and visual identity | Not answered |
| Lesson-player UX outside education | Not answered |
| Onboarding to first value | Not answered |
Relevant because we have already shipped XP, streaks, badges and a confidence rating.
From Deci, Koestner and Ryan's meta-analysis in Psychological Bulletin, every figure verbatim from the paper. Tangible rewards measurably undermine intrinsic motivation. But the finding that matters most for us is about praise, the mechanism designers assume is safe.
The authors put it plainly: "although verbal rewards enhance intrinsic motivation for college students, they do not do so for children." Same mechanic, opposite sign, decided entirely by wording.
Deci, Koestner & Ryan (1999), Psychological Bulletin 125(6), 627-668. Disputed on interpretation by Cameron & Pierce and by Eisenberger et al.
Sailer and Homner (2020), Educational Psychology Review, verbatim from the abstract:
"significant small effects of gamification on cognitive (g = 0.49), motivational (g = 0.36), and behavioral learning outcomes (g = 0.25). Whereas the effect of gamification on cognitive learning outcomes was stable in a subsplit analysis of studies employing high methodological rigor, effects on motivational and behavioral outcomes were less stable."
Read that carefully. The learning benefit holds up. The motivation and behaviour benefits do not. The thing an XP and streak layer exists to do, bring a child back tomorrow, is the least-evidenced claim in the field. Later work still calls intrinsic-motivation findings inconsistent and flags novelty-effect decay.
The same meta-analysis found game fiction moderates behavioural outcomes, and competitive-plus-collaborative designs beat no-social-interaction designs. But motivational outcomes showed no significant difference in the main analysis (p = .28), appearing only in the high-rigour subsplit, and the behavioural result rests on nine studies split three ways.
Mild support for our quest-map framing. It does not establish that a quest map beats a plain XP counter. Not worth overselling.
Two precedents worth copying almost directly.
Three things, verbatim from Khan's own documentation:
That last one is the one to steal. Khan lets a parent be a monitor rather than a teacher, and says so in plain language. Parent-side assignment creation only arrived in 2025, so this is a current deliberate choice rather than legacy copy.
One more: clicking the person icon beside a child's name drops the parent into that child's own Learner Home. A real preview of the child's view, not a mock-up of it.
Khan Academy Help Center. Direct fetch was blocked (HTTP 403), so confirmation came via search extraction. Nobody read these in a browser session.
From IXL's parent analytics guide, under the heading "How can I help my child with areas they're struggling with?":
"It pinpoints the concepts your child is struggling with, and you can even see the exact questions your child received and missed!"
Then three concrete moves, and critically a stopping point:
The report tells the parent what to do and when to stop. Ours shows that Ruby felt "a little unsure" and leaves her there.
Ranked by leverage, not by effort.
The lesson page currently says "Quest complete! Nicely done, Ruby!" That is textbook controlling praise, the -0.44 case. Informational praise is +0.66. Name what she actually did: the skill, the step, what changed since last time. Best evidence-to-effort ratio on this list, and it is a copy change.
Khan explicitly lets a parent skip the teacher role. We gate everything behind generate-then-approve, so a parent who never touches admin sees nothing at all. Same wound as the draft dead end in the flow map.
Name the specific questions, give two or three concrete moves, define when to stop. IXL's shape, applied to our confidence data.
Keep it, it is cheap and harmless, but the evidence says it will not carry the load we are implicitly asking of it. A streak also punishes absence, and a homeschool term has holidays and sick days built into it.
Learn it from Prodigy rather than repeating it. Relevant the moment this venture has a paid tier.
Praise does not help children, and reverses when it is controlling. But we already collect something most products do not: a self-reported confidence rating, one to five, at the exact moment of completion.
That gives us an informational completion moment with no evaluative praise in it at all. "You said you felt unsure about this one, and you got every question right." That is information about the child's own calibration, not a judgement handed down. It is the +0.66 shape rather than the -0.44 shape, and we are already capturing the data needed to say it.
Fifteen of twenty five verified claims were refuted. The casualties cluster exactly where a product team most likes to cite.
On a normal search-and-summarise, most of those would have reached you as fact. That is what the expensive verification pass actually bought.
A second pass, one question at a time, on a cheaper model split for the high-volume work: