Design research ยท 19 July 2026

What the evidence actually says about rewarding children.

A deep research pass with adversarial verification. Every claim here survived three independent agents trying to knock it down. Fifteen claims did not survive, and those are listed too.

QuestionStatus
Parent dashboard and progress reportingAnswered
Child motivation systems, with critiqueAnswered
Dual parent/child mode patternsNot answered
Typography and visual identityNot answered
Lesson-player UX outside educationNot answered
Onboarding to first valueNot answered

The finding that changes our build

Relevant because we have already shipped XP, streaks, badges and a confidence rating.

Praise is not neutral High confidence

From Deci, Koestner and Ryan's meta-analysis in Psychological Bulletin, every figure verbatim from the paper. Tangible rewards measurably undermine intrinsic motivation. But the finding that matters most for us is about praise, the mechanism designers assume is safe.

-0.34Expected tangible rewards, effect on intrinsic motivation (92 studies)
0.11Verbal praise, children specifically. Not statistically significant
-0.44Controlling praise. "Well done, you did what I wanted"
+0.66Informational praise. Specific about what the child actually did

The authors put it plainly: "although verbal rewards enhance intrinsic motivation for college students, they do not do so for children." Same mechanic, opposite sign, decided entirely by wording.

Deci, Koestner & Ryan (1999), Psychological Bulletin 125(6), 627-668. Disputed on interpretation by Cameron & Pierce and by Eisenberger et al.

Gamification's motivational effect does not survive scrutiny High confidence

Sailer and Homner (2020), Educational Psychology Review, verbatim from the abstract:

"significant small effects of gamification on cognitive (g = 0.49), motivational (g = 0.36), and behavioral learning outcomes (g = 0.25). Whereas the effect of gamification on cognitive learning outcomes was stable in a subsplit analysis of studies employing high methodological rigor, effects on motivational and behavioral outcomes were less stable."

Read that carefully. The learning benefit holds up. The motivation and behaviour benefits do not. The thing an XP and streak layer exists to do, bring a child back tomorrow, is the least-evidenced claim in the field. Later work still calls intrinsic-motivation findings inconsistent and flags novelty-effect decay.

Narrative and cooperation, weakly Medium confidence

The same meta-analysis found game fiction moderates behavioural outcomes, and competitive-plus-collaborative designs beat no-social-interaction designs. But motivational outcomes showed no significant difference in the main analysis (p = .28), appearing only in the high-rigour subsplit, and the behavioural result rests on nine studies split three ways.

Mild support for our quest-map framing. It does not establish that a quest map beats a plain XP counter. Not worth overselling.

Prodigy, the documented anti-pattern High confidence

Parent dashboards

Two precedents worth copying almost directly.

Khan Academy: let a parent opt out of being a teacher High confidence

Three things, verbatim from Khan's own documentation:

That last one is the one to steal. Khan lets a parent be a monitor rather than a teacher, and says so in plain language. Parent-side assignment creation only arrived in 2025, so this is a current deliberate choice rather than legacy copy.

One more: clicking the person icon beside a child's name drops the parent into that child's own Learner Home. A real preview of the child's view, not a mock-up of it.

Khan Academy Help Center. Direct fetch was blocked (HTTP 403), so confirmation came via search extraction. Nobody read these in a browser session.

IXL Trouble Spots: never stop at a number High confidence

From IXL's parent analytics guide, under the heading "How can I help my child with areas they're struggling with?":

"It pinpoints the concepts your child is struggling with, and you can even see the exact questions your child received and missed!"

Then three concrete moves, and critically a stopping point:

The report tells the parent what to do and when to stop. Ours shows that Ruby felt "a little unsure" and leaves her there.

What this changes in our build

Ranked by leverage, not by effort.

  1. Rewrite the celebration copy

    The lesson page currently says "Quest complete! Nicely done, Ruby!" That is textbook controlling praise, the -0.44 case. Informational praise is +0.66. Name what she actually did: the skill, the step, what changed since last time. Best evidence-to-effort ratio on this list, and it is a copy change.

  2. Give parents a monitoring-only path

    Khan explicitly lets a parent skip the teacher role. We gate everything behind generate-then-approve, so a parent who never touches admin sees nothing at all. Same wound as the draft dead end in the flow map.

  3. Make work review prescriptive

    Name the specific questions, give two or three concrete moves, define when to stop. IXL's shape, applied to our confidence data.

  4. Stop treating the streak as the engagement engine

    Keep it, it is cheap and harmless, but the evidence says it will not carry the load we are implicitly asking of it. A streak also punishes absence, and a homeschool term has holidays and sick days built into it.

  5. Never render paid or tiered status inside the child's space

    Learn it from Prodigy rather than repeating it. Relevant the moment this venture has a paid tier.

The confidence rating is quietly our best asset

Praise does not help children, and reverses when it is controlling. But we already collect something most products do not: a self-reported confidence rating, one to five, at the exact moment of completion.

That gives us an informational completion moment with no evaluative praise in it at all. "You said you felt unsure about this one, and you got every question right." That is information about the child's own calibration, not a judgement handed down. It is the +0.66 shape rather than the -0.44 shape, and we are already capturing the data needed to say it.

What got killed

Fifteen of twenty five verified claims were refuted. The casualties cluster exactly where a product team most likes to cite.

On a normal search-and-summarise, most of those would have reached you as fact. That is what the expensive verification pass actually bought.

Limits

Next

A second pass, one question at a time, on a cheaper model split for the high-volume work: