Key takeaways
- Same AI model, wildly different outcomes: productive struggle ranged from 0.1% to 40.4% across three deployments, tracking teacher, parent, and peer involvement, not model quality.
- This behaviour predates generative AI — a pre-LLM tool with 4M+ users showed the same shortcut-seeking behaviour.
- Parents underestimate their children’s use of AI for schoolwork.
- CoLearn predicts full in-school adoption, with high involvement across the board, could push the rate above 50% — not yet observed.
- Implication for builders: “who’s in the loop” should be a core design variable, and the AI accuracy bar rises fast at institutional scale.
The fact that students learn more durably when they wrestle with a problem at the edge of their ability before receiving help, rather than being handed a solution path immediately, is well established (Roediger & Karpicke, 2006; Hiebert & Grouws, 2007; Kapur & Roll, 2023; Oakley et al., 2025).
The global rise of generative AI complicates this: because it is so effective at producing correct answers in seconds, its reach has already had documented negative implications for learning worldwide (Burns et al., 2026). However, it is clear to us that generative AI also holds huge potential in the other direction — to widen access to high-quality tutoring for the masses, offering personalized help that still promotes the productive struggle genuine learning requires.
Since October 2025, CoLearn has released its first iteration of an AI Tutor — powered by Gemini’s 2.5 series and built to require students to submit their own written work before receiving help — into three different settings:
- October 2025 — as a homework and exam-prep service for students in CoLearn’s cohort-based live classes, promoted by teachers in the classroom and communicated to parents, but positioned as an additional service outside the classroom itself.
- December 2025 — as a standalone app available internationally, aimed at helping students with homework and quizzes, built around one core principle: no effort, no help.
- April 2026 — piloted directly inside the live-class curriculum as part of a weekly Practice session, complete with a leaderboard and progress reports sent to parents.
We didn’t go in with a fixed idea about which setting would work best — we just wanted to see how well each one got students to do real, handwritten work before turning to AI for help. Over the next few weeks, we looked at both the usage data and what we learned firsthand: sitting down with parents and students in person, and hearing from users around the world who were willing to share their feedback.
Results
The clearest signal came from a simple, hard-to-game proxy: the share of students still submitting handwritten work — evidence of an actual attempt — four weeks after starting to use the AI tutor, rather than work that reads as copied straight out of an AI-generated response. Across three deployments built on the same underlying tutoring capability, that rate ranged from effectively zero to four in ten.
AI Tutor International — a purely self-serve product with no teacher, parent, or peer involvement built in — held productive struggle at 0.1%. AI Tutor Indonesia, structurally similar but with some low-touch teacher involvement, reached 5.2%. Practice, which layers in high teacher involvement plus some parent and peer involvement, reached 40.4%. For a full in-school deployment — complete teacher integration, embedded in the classroom rather than offered as a take-home supplement, with correspondingly high parent and peer involvement too — CoLearn predicts the rate would clear 50%, though that configuration hasn’t yet been observed at scale.
The pattern that stands out isn’t that a better model changes the outcome — all three products draw on comparable underlying tutoring capability — it’s that the social scaffolding wrapped around the model does. That maps fairly directly onto older ideas in learning theory: a Vygotskian zone of proximal development is, by definition, mediated by a more knowledgeable other, not by a static resource a student consults alone. A teacher who can see whether work was actually attempted, a parent who asks about homework at dinner, or classmates visibly doing their own problem sets all appear to function as exactly that kind of mediation — even when the underlying AI system is identical.
Why "instant answer" is the default expectation
Demand for AI-aided help with schoolwork is undeniably strong — but not necessarily for reasons that improve learning. Install and usage numbers for a self-serve tutor climb quickly once it’s available; the harder question is what students actually want from it once they’re in.
Sitting down with students from primary through high school age, alongside their parents, one thing became obvious: most students in middle and high school already have access to publicly available generative AI, and most expect it to produce an instant answer within seconds. Unless specifically required to do otherwise, they won’t opt into a tool that asks them to try first before getting help, or one designed to give scaffolded guidance that builds long-term understanding. The short-term expectation of getting today’s homework done wins out over the long-term benefit of actually learning the material.
The behavior predates the technology
Before building any generative-AI product, CoLearn ran a step-by-step video tool that reached more than 4 million monthly active users in Indonesia. Each video walked through a problem’s solution incrementally, in contrast to platforms that simply returned a final text answer. The instructional design was deliberately scaffolded. It didn’t matter: most students used the path of least resistance anyway, scrolling directly to the end of the video to read off the answer rather than following the steps.
That baseline is worth sitting with before assigning today’s “students using ChatGPT to skip the thinking” narrative entirely to large language models. The tendency to route around effortful learning when a shortcut is available showed up at scale in a product that had no LLM in it at all. Whatever is driving answer-seeking behavior, it isn’t solely a property of how capable or how conversational the underlying model is.
What this data can and can't support
This comparison is observational, not a randomized trial, and the three products were built into different environments to begin with. What strengthens our observation, though, is the pre-LLM baseline described below: the same struggle-avoidance behavior showed up in a product with no involvement mechanism and no LLM at all, at a scale too large to be an artifact of one tool’s design. That’s consistent with — though it doesn’t prove — an explanation centered on the surrounding human system rather than the AI system’s own pedagogy.
No adults in the room
Just as striking was how many parents — especially of middle and high schoolers — are unaware of how often, or how, their children are already using generative AI for schoolwork. When this came up, the most common response was to reach for their own childhood as a reference point: they had to write out their answers and struggle through them, so surely their children are doing the same. By contrast, the students we met who were genuinely regular, productive users of the AI tutor tended to have parents who were closely aware of their progress — or, in rarer cases, were unusually self-motivated students chasing something beyond the curriculum, like a national olympiad.
One child told us he only used AI “sometimes” when a problem was too hard. When we asked the child to put a number on it, he said eight out of ten. The mother hadn’t known.
That gap points to a pattern that’s easy to miss when an educational product is designed mainly around the relationship between student, teacher, and content: parents’ awareness and involvement is itself a significant driver of whether AI gets used productively. It matters, too, that most parent–student conversations about schoolwork focus on the final outcome — the grade, the completed assignment — and rarely touch the process behind it. In a system that rewards the outcome and rarely inspects the process, it isn’t hard to see why students would reach for the shortcut that produces an outcome their parents will accept, whether that’s an AI-generated answer or, in more cases than anyone would like, copying on an exam.
Where do we go from here
If this pattern holds up under more rigorous testing, it argues for treating “who is in the loop” — teacher, parent, peer — as a first-order design variable in AI tutoring, on par with model accuracy and prompting strategy, rather than a secondary implementation detail. The parent-awareness gap suggests that loop needs to be actively designed into a product, not assumed: a leaderboard or progress report that only reaches a parent who already checks in is not the same as one that prompts a parent who doesn’t.
That leaves a fairly concrete agenda. What’s the minimum effective dose of teacher involvement needed to move the needle? Can peer effects be deliberately engineered as a product feature — visible shared progress, cohort framing — rather than left to emerge incidentally from a classroom setting? And does a proxy like handwritten-submission rate actually track learning gains, or does it simply track effort that may or may not translate into understanding? None of these questions can be answered from one organization’s product telemetry alone — they need study designs built specifically to isolate the mechanism, ideally across multiple providers, products, and contexts.
References
Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
Hiebert, J., & Grouws, D. A. (2007). The effects of classroom mathematics teaching on students’ learning. In F. K. Lester Jr. (Ed.), Second handbook of research on mathematics teaching and learning (pp. 371–404). Information Age Publishing / National Council of Teachers of Mathematics.
Kapur, M., Saba, J. & Roll, I. Prior math achievement and inventive production predict learning from productive failure. npj Sci. Learn. 8, 15 (2023). https://doi.org/10.1038/s41539-023-00165-y
Oakley, B., Johnston, M., Chen, K., Jung, E., & Sejnowski, T. (2025). The memory paradox: Why our brains need knowledge in an age of AI. SSRN / arXiv:2506.11015. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5250447
Burns, M., Winthrop, R., Luther, N., Venetis, E., & Karim, R. (2026, January 14). A new direction for students in an AI world: Prosper, prepare, protect. Brookings Institution, Center for Universal Education. https://www.brookings.edu/articles/a-new-direction-for-students-in-an-ai-world-prosper-prepare-protect/
