Within Learn LM Maths
Why the AI tutor still needed humans
The trial's strongest lesson may be that AI tutoring looked promising because expert humans constrained and checked the system.
On this page
- What human supervisors reviewed before students saw messages
- The risks oversight was designed to reduce
- What this means for scaling AI tutoring safely
Page outline Jump by section
Introduction
The most important lesson from the Eedi–LearnLM maths tutoring trial may not be that an AI tutor performed well. It may be that the system performed well because expert humans remained deeply involved.
In the UK study, students did not interact directly with an unrestricted AI model. Instead, LearnLM generated draft tutoring messages which were reviewed by expert tutors before reaching students. Human supervisors could approve, edit, or completely rewrite every response. The trial therefore tested a human-supervised tutoring system rather than autonomous AI teaching. That distinction matters because many of the reported gains in reliability, safety and educational quality may have depended on the presence of those human checks.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
For advocates of AI-enabled educational abundance, the trial is encouraging. It suggests that AI may help deliver more personalised support to many more learners. But it also points to a harder implementation question: if high-quality outcomes depend on expert oversight, how much human judgement is still required before AI tutoring can be safely expanded to millions of students?
Why the AI tutor still needed humans
One common misconception about the Eedi trial is that it demonstrated AI replacing tutors. The actual design was closer to a partnership.
LearnLM generated responses during mathematics tutoring sessions, but supervising tutors retained final authority. Researchers instructed tutors to revise every draft until they would personally be comfortable sending it themselves. In practice, tutors could approve messages unchanged, make small edits, substantially rewrite them, or discard them entirely.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
This arrangement reflected a broader reality about educational AI. Mathematics tutoring is not simply about producing correct answers. Tutors must diagnose misconceptions, judge how much help to provide, decide when to ask a question instead of giving an explanation, and recognise when a student is confused, disengaged or heading down the wrong path.
Current language models can often solve mathematical problems. They are less consistently reliable at making pedagogical decisions about how a student should learn those problems. The human supervisors acted as a safeguard against mistakes in that judgement.[Dan Meyer]danmeyer.substack.comDan MeyerResearch Review: AI+Human Tutors Match Quality of…December 17, 2025 — The researchers embedded three kinds of support inside…
The design also acknowledged a basic asymmetry of risk. A strong tutoring message can help learning. A poor tutoring message can reinforce misunderstandings, provide shortcuts that bypass reasoning, or leave students with false confidence. Because educational interactions accumulate over time, even small errors can matter.
What human supervisors reviewed before students saw messages
The supervising tutors were not merely scanning for offensive content. Their role was much broader and more educationally specific.
Before messages reached students, tutors could evaluate whether LearnLM’s suggestions:
- Were mathematically correct.
- Matched the student’s apparent misconception.
- Encouraged reasoning rather than answer-copying.
- Asked useful follow-up questions.
- Maintained an appropriate difficulty level.
- Used clear and age-appropriate language.
- Avoided revealing solutions too quickly.
- Fit accepted teaching practice.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…[Dan Meyer]danmeyer.substack.comDan MeyerResearch Review: AI+Human Tutors Match Quality of…December 17, 2025 — The researchers embedded three kinds of support inside…
The trial results suggest that LearnLM often produced acceptable drafts. Researchers reported that supervising tutors approved roughly three-quarters of AI-generated messages with either no edits or only minimal changes. Different reports of the study cite figures ranging from about 74% to over 80%, depending on the analysis and threshold used.[Eedi Labs]eedi-labs-67qwnm2lxu4blqi5.webflow.ioEedi LabsNew Exploratory Research from Eedi and Google…The trial, conducted with 165 students and 17 tutors, was a careful test of saf…
That finding is important because it shows the human reviewers were not rewriting everything. The AI appeared capable of producing many useful tutoring prompts on its own. Yet the remaining quarter of messages still required meaningful intervention, which illustrates why supervision remained valuable.
Several tutors reported that LearnLM was particularly strong at drafting Socratic questions — prompts that encourage students to explain their thinking rather than simply receive answers. Some even reported learning new questioning approaches from the model. But these reports emerged within a framework where humans could still filter and refine the AI’s suggestions.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
The risks oversight was designed to reduce
Educational AI systems face a different risk profile from many consumer chatbots because they are interacting with learners who may not recognise mistakes.
Preventing factual and mathematical errors
The most obvious concern is incorrect information.
A student working through algebra or geometry may not know whether an explanation is right or wrong. An error delivered confidently by a tutoring system can therefore become a learned misconception rather than a simple mistake.
The Eedi trial included auditing of AI-generated interactions and reported extremely low rates of factual inaccuracies. Human review was part of the mechanism intended to achieve that result.[Eedi Labs]eedi-labs-67qwnm2lxu4blqi5.webflow.ioEedi LabsNew Exploratory Research from Eedi and Google…The trial, conducted with 165 students and 17 tutors, was a careful test of saf…
Preventing over-helping
A less obvious risk is pedagogical failure.
A model can help a student solve a problem while simultaneously undermining learning. For example, it may reveal too much information too quickly, provide procedural steps without understanding, or guide students directly to answers.
Human tutors regularly make subtle decisions about when not to help. Educational researchers have long noted that productive struggle can be important for learning. The Eedi supervisors could intervene when AI suggestions became too leading or insufficiently challenging.[Eedi Newsletter]eedi.substack.comEedi Newsletter AI in Education #7: Five Things I Learned from Bibi GrootLearnLM with a human in the loop approving every message before it reaches the student. It is not definitive proof that AI tutoring works…
Preventing unsafe or inappropriate outputs
Another goal was basic safety.
Public concern about educational chatbots often focuses on hallucinations, inappropriate content or unpredictable behaviour. The study’s safety audits reported no harmful content reaching students during the trial. While LearnLM itself was specifically adapted for educational use, the presence of human review created an additional protective layer.[Eedi Labs]eedi-labs-67qwnm2lxu4blqi5.webflow.ioEedi LabsNew Exploratory Research from Eedi and Google…The trial, conducted with 165 students and 17 tutors, was a careful test of saf…
Preserving accountability
Oversight also solved an institutional problem.
Schools, teachers and parents generally know how to hold human educators accountable. Accountability becomes murkier when decisions emerge from probabilistic AI systems. By keeping expert tutors responsible for final messages, the trial maintained a clear chain of responsibility.
That may prove important beyond mathematics. If AI becomes involved in educational guidance, assessment, feedback or mentoring, institutions will likely want identifiable humans who can justify and defend important decisions.
What the trial suggests about scaling AI tutoring
The trial’s most optimistic interpretation is that human supervision and AI generation can complement one another rather than compete.
Researchers reported that tutors often felt able to manage larger workloads when supported by LearnLM. Informal follow-up tests suggested that supervised tutors could handle more simultaneous conversations than tutors working entirely manually. The AI handled much of the drafting work, while humans focused attention on quality control and intervention.[Google Cloud Storage]storage.googleapis.comGoogle Cloud StorageAI tutoring can safely and effectively support studentsNovember 10, 2025 — 11 Nov 2025 — Overall, the design of this…
That matters for the larger AI bloom argument around education and cognitive empowerment.
One of the strongest cases for educational AI is not that every learner receives a perfect artificial teacher. It is that scarce educational expertise could be amplified. A skilled tutor might supervise support for many more students than would otherwise be possible, potentially expanding access to personalised learning beyond affluent households that can currently afford one-to-one tutoring.
In that vision, AI functions less as a replacement for educational expertise and more as a force multiplier for it.
Yet the trial also exposes a practical bottleneck. If every message requires human approval, scaling remains constrained by the availability of expert reviewers. The system becomes more efficient than traditional tutoring, but not infinitely scalable.
From human-in-the-loop to human-on-the-loop
One of the most interesting debates emerging from the Eedi work concerns what happens next.
The trial used a classic “human-in-the-loop” design. Every student-facing message passed through human review. Some researchers and practitioners have suggested that future systems may move toward a “human-on-the-loop” model instead, where AI handles most interactions autonomously while humans monitor conversations and intervene when warning signs appear.[Eedi Newsletter]eedi.substack.comEedi Newsletter AI in Education #7: Five Things I Learned from Bibi GrootLearnLM with a human in the loop approving every message before it reaches the student. It is not definitive proof that AI tutoring works…
The attraction is obvious. Continuous review of every message is expensive. Monitoring exceptions is cheaper.
But moving to lighter supervision would change the evidence base. The trial’s positive results apply to a system with intensive oversight. They do not automatically demonstrate that an autonomous tutor would achieve the same safety or educational outcomes.
This is a recurring pattern across AI deployment. Early successes often emerge under carefully controlled conditions with expert operators, extensive monitoring and constrained environments. The challenge comes when organisations try to preserve those gains while reducing human involvement.
The Eedi study therefore provides evidence not only about AI tutoring but also about a broader governance question: how much human oversight is necessary before society can trust AI systems in high-stakes domains?
What this means for AI abundance and human flourishing
The trial offers a small but concrete glimpse of a larger possibility behind AI bloom.
If advanced AI can help skilled educators reach more learners without significantly reducing quality, educational support could become far more abundant. Personalised tutoring has long been one of education’s most effective interventions, yet it remains scarce because expert human attention is expensive. AI systems may help relax that constraint.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
At the same time, the study points away from simplistic automation narratives.
The strongest result was not “AI replaces teachers”. It was that a carefully constrained system, shaped by pedagogical design and checked by expert humans, appeared capable of supporting learning at high quality. The human supervisors were not a temporary inconvenience added before full automation. They were part of the intervention itself.
That distinction matters for long-term discussions about superintelligence, abundance and human flourishing. Many optimistic visions assume advanced AI will eventually make expertise widely available. The Eedi trial suggests that, at least in the near term, the path may involve hybrid systems where AI expands human capability rather than eliminating human judgement.
For education, that may be the more important lesson. The future suggested by the trial is not one in which machines simply take over teaching. It is one in which human expertise becomes more scalable because intelligent systems help deliver it more widely, while people remain responsible for deciding what good teaching actually looks like.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…[Eedi Newsletter]eedi.substack.comEedi Newsletter AI in Education #7: Five Things I Learned from Bibi GrootLearnLM with a human in the loop approving every message before it reaches the student. It is not definitive proof that AI tutoring works…
Endnotes
1.
Source: arxiv.org
Link:https://arxiv.org/abs/2512.23633
Source snippet
AI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025...
Published: December 29, 2025
2.
Source: eedi.com
Link:https://www.eedi.com/learn
Source snippet
LearnOur expert human tutors review, edit, and approve each response from the AI tutor before it is sent on to the student; Once the misc...
3.
Source: eedi.com
Link:https://www.eedi.com/news/just-launched—our-second-ai-tutor-rct
Source snippet
instances of harmful content, and our Bayesian analysis...Read more...
4.
Source: arxiv.org
Link:https://arxiv.org/html/2512.23633v1
Source snippet
AI tutoring can safely and effectively support students25 Nov 2025 — Overall, the design of this exploratory RCT allowed us to rapid...
5.
Source: eedi.com
Link:https://www.eedi.com/news/new-uk-study-finds-students-using-eedi-gain-2-4-months-of-additional-maths-progress
Source snippet
New UK Study Finds Students Using Eedi Gain 2–4...7 Nov 2025 — Independent randomised controlled trial across 20 schools and 3,000 stude...
6.
Source: danmeyer.substack.com
Link:https://danmeyer.substack.com/p/research-review-aihuman-tutors-match
Source snippet
Dan MeyerResearch Review: AI+Human Tutors Match Quality of...December 17, 2025 — The researchers embedded three kinds of support inside...
Published: December 17, 2025
7.
Source: eedi-labs-67qwnm2lxu4blqi5.webflow.io
Link:https://eedi-labs-67qwnm2lxu4blqi5.webflow.io/news/new-exploratory-research-from-eedi-and-google-deepmind-reveals-human-in-the-loop-ai-tutoring-outperforms-human-only-support
Source snippet
Eedi LabsNew Exploratory Research from Eedi and Google...The trial, conducted with 165 students and 17 tutors, was a careful test of saf...
8.
Source: eedi.substack.com
Title: Eedi Newsletter AI in Education #7: Five Things I Learned from Bibi Groot
Link:https://eedi.substack.com/p/ai-in-education-7-five-things-i-learned
Source snippet
LearnLM with a human in the loop approving every message before it reaches the student. It is not definitive proof that AI tutoring works...
9.
Source: storage.googleapis.com
Link:https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_nov25.pdf
Source snippet
Google Cloud StorageAI tutoring can safely and effectively support studentsNovember 10, 2025 — 11 Nov 2025 — Overall, the design of this...
Published: November 10, 2025
Additional References
10.
Source: linkedin.com
Link:https://www.linkedin.com/posts/learning-agency_ai-tutors-with-a-little-human-help-offer-activity-7402719664608288768-Rehs
Source snippet
AI Tutors Outperform Humans with Human SupervisionAs The 74 Media reports, “students using the supervised AI tutor performed slightly bet...
11.
Source: linkedin.com
Link:https://www.linkedin.com/pulse/why-deeper-questioning-led-better-learning-unpacking-our-bibi-groot-6yore
Source snippet
Why deeper questioning led to better learningThe study compared three conditions: static hints, human-only tutoring, and LearnLM (supervi...
12.
Source: learning-engineering-virtual-institute.org
Link:https://learning-engineering-virtual-institute.org/ai-tutoring-can-safely-and-effectively-support-students-an-exploratory-rct-in-uk-classrooms/
Source snippet
Learning Engineering Virtual InstituteAI Tutoring Can Safely And Effectively Support Students28 Nov 2025 — LearnLM proved to be a reliabl...
13.
Source: linkedin.com
Link:https://www.linkedin.com/posts/simon-woodhead_edtech-aied-edm-activity-7393960569541652481-kR0c
Source snippet
Eedi and Google DeepMind: Human-Supervised LearnLM...Built-in accountability and recourse: Human tutors were ultimately responsible for...
14.
Source: edtechinnovationhub.com
Title: eedi and google deepmind begin second ai tutoring trial across 1525 uk students
Link:https://www.edtechinnovationhub.com/news/eedi-and-google-deepmind-begin-second-ai-tutoring-trial-across-1525-uk-students
Source snippet
Eedi and DeepMind trial AI tutor with 1525 UK students5 May 2026 — Supervising tutors approved 74.4 percent of AI-drafted messages withou...
Published: May 2026
15.
Source: the74million.org
Title: ai tutors with a little human help offer reliable instruction study finds
Link:https://www.the74million.org/article/ai-tutors-with-a-little-human-help-offer-reliable-instruction-study-finds/
Source snippet
AI Tutors, With a Little Human Help, Offer 'Reliable'...3 Dec 2025 — New UK research finds 'extremely personalized' AI math tutors don't...
16.
Source: brookings.edu
Title: what the research shows about generative ai in tutoring
Link:https://www.brookings.edu/articles/what-the-research-shows-about-generative-ai-in-tutoring/
Source snippet
27 Jan 2026 — U.K.: (LearnLM Team, Google & Eedi, 2025), • An exploratory RCT with 165 students across five UK secondary schools focused...
17.
Source: finance.yahoo.com
Title: exploratory research eedi google deepmind 090000225
Link:https://finance.yahoo.com/news/exploratory-research-eedi-google-deepmind-090000225.html
Source snippet
Exploratory Research From Eedi and Google...11 Nov 2025 — A new exploratory study from Eedi, an evidence-based provider of AI-powered le...
18.
Source: eedi.substack.com
Title: ai empowers the teacher so they can
Link:https://eedi.substack.com/p/ai-empowers-the-teacher-so-they-can
Source snippet
empowers the teacher so they can help more childrenToday marks an exciting day in the history of Eedi. The results of our year-long work...
19.
Source: erctpapers.com
Link:https://erctpapers.com/papers/144-team-ai-tutoring-can-safely-and-effectively-support-students-an-exploratory-rct-in-uk.html
Source snippet
AI tutoring can safely and effectively support students - ERCT29 Dec 2025 — AI tutoring can safely and effectively support students: An e...
Topic Tree



