Within Learn LM Maths
How strong is the Learn LM evidence?
The LearnLM result is encouraging, but its small UK sample and short sessions make it evidence to test further, not proof at scale.
On this page
- What the UK randomised trial measured
- Why the 5.5 point transfer result matters
- What remains uncertain after a small short study
Page outline Jump by section
Introduction
The 165-student Eedi trial is one of the most interesting early studies in AI-assisted education because it tested something harder than whether a chatbot can answer maths questions. It tested whether a carefully designed AI tutoring system could help students learn mathematics well enough to apply ideas to new problems later. The result was encouraging: students receiving human-supervised LearnLM tutoring performed at least as well as students working with human tutors, and in one transfer measure they did slightly better.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
But the study does not prove that AI tutors can transform education at national scale, replace teachers, or reliably deliver large learning gains across subjects and age groups. What it really provides is something narrower and arguably more important: early evidence that a pedagogically tuned AI system can participate in genuine tutoring interactions without obviously degrading learning outcomes. In the context of larger debates about AI abundance and cognitive empowerment, that is a meaningful signal. It is not yet a definitive answer.
What the UK randomised trial measured
The study involved 165 students aged roughly 13–15 across five UK secondary schools using the Eedi mathematics platform. Students were randomly assigned to different forms of support after encountering difficulties with maths questions. The trial ran during May and June 2025.[Google Cloud Storage]storage.googleapis.comlearn LM nov25Each student and each…
The key comparison was not simply “AI versus no AI”. Researchers compared several forms of help:
- Static pre-written hints already available on the platform.
- One-to-one human tutoring through chat.
- LearnLM-generated tutoring responses reviewed by expert human tutors before being sent to students.[Google Cloud Storage]storage.googleapis.comlearn LM nov25Each student and each…
That design matters because many public discussions of AI tutoring compare highly constrained educational systems with unrestricted consumer chatbots. The Eedi trial did not do that. Students interacted with an AI operating inside a structured educational environment, focused on specific mathematical misconceptions and supervised by trained tutors.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
The trial therefore provides evidence about a particular model of AI-assisted tutoring: tightly bounded, pedagogically tuned, and human-supervised.
Why the 5.5-point transfer result matters
The most discussed finding was not immediate question completion. It was a transfer measure.
Students receiving LearnLM-supported tutoring solved novel subsequent problems at a rate of 66.2%, compared with 60.7% for students receiving human-only tutoring. That difference amounted to 5.5 percentage points.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
The reason researchers pay attention to transfer is that education is ultimately about applying knowledge beyond the exact exercise that was just practised. A student who can answer only the question they have just seen has not necessarily learned much. A student who can apply the underlying concept in a new context may have developed genuine understanding.
In AI education research, this distinction is particularly important. Large language models can often help students arrive at correct answers. The harder question is whether they help students build durable mental models.
The transfer result therefore carries more weight than a simple increase in completion rates would. If real, it suggests the AI system may have encouraged forms of reasoning that generalised beyond the immediate interaction. Tutors interviewed after the study reported that LearnLM often generated useful Socratic-style questions that prompted reflection rather than immediately revealing answers. Some tutors even reported learning new tutoring techniques from the model itself.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
For advocates of a broader AI bloom future, this is the genuinely interesting part. The long-term promise of educational AI is not merely cheaper homework assistance. It is the possibility of making high-quality intellectual guidance much more widely available. Evidence of transfer learning is closer to that goal than evidence of answer generation.
What the study can reasonably support
Several conclusions appear justified by the evidence.
First, the study suggests that pedagogically tuned AI can contribute meaningfully to tutoring interactions. The system was not merely generating fluent text. Students learning with AI-supported tutoring performed at least as well as those receiving human-only tutoring on the measured outcomes.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
Second, it suggests that carefully constrained AI systems may be safer and more reliable than many critics assume. Supervising tutors approved roughly three quarters of LearnLM-generated messages with either no edits or only trivial edits. That does not mean the system was flawless, but it indicates that a large share of its responses were considered suitable by experienced educators.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
Third, the study provides evidence that educational AI may be most effective when designed around learning science rather than around general conversational ability. LearnLM was explicitly trained to support pedagogical interaction, questioning, scaffolding and student reasoning. The trial therefore supports a broader lesson: educational outcomes may depend heavily on system design choices rather than simply on model size or raw capability.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
Finally, the trial strengthens the case for further experimentation. Randomised controlled trials remain one of the strongest methods available for measuring educational interventions. Even a small exploratory RCT generally provides more useful evidence than anecdotal classroom reports or marketing claims.[scale.stanford.edu]scale.stanford.eduRandomized Controlled TrialSTART HERE for the most rigorous research on…
What remains uncertain after a small short study
The study’s limitations are at least as important as its positive findings.
The sample was small
A total of 165 students is a respectable pilot but a modest sample for making strong claims about national education systems. Small studies can produce results that later shrink, disappear, or reverse when tested across larger populations.
The difference between 66.2% and 60.7% may represent a genuine effect. It may also prove sensitive to population differences, school environments, implementation quality, or statistical noise. Larger replication studies are needed before treating the result as settled.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
The intervention was short
The trial ran across a limited period in May and June 2025.[Google Cloud Storage]storage.googleapis.comlearn LM nov25Each student and each…
That means it primarily measured short-term learning outcomes rather than long-term educational development.
Important unanswered questions include:
- Do gains persist months later?
- Do students retain concepts better?
- Does AI tutoring improve examination performance?
- Does it affect confidence, motivation, or study habits?
- Does repeated exposure change how students approach problem-solving?
A system can improve immediate learning while producing weaker long-term effects. Equally, some benefits may only become visible after extended use.
Human supervision did a great deal of work
One common misunderstanding is that the trial proved autonomous AI tutors outperform humans.
That is not what happened.
Human experts reviewed LearnLM’s outputs before students saw them. The trial therefore tested a human-AI team rather than an independent AI tutor.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
This distinction matters because scaling to millions of students may require different supervision arrangements. A model that performs well with expert oversight may not behave identically when deployed with minimal monitoring.
The result therefore supports claims about supervised educational AI more strongly than claims about fully autonomous tutoring systems.
The setting was unusually structured
The Eedi platform already centres on diagnosing mathematical misconceptions and delivering targeted interventions. Students were not engaging in unrestricted conversations across arbitrary topics.[Google Cloud Storage]storage.googleapis.comlearn LM nov25Each student and each…
That structure likely contributed to the positive outcomes.
A constrained system operating within a well-defined curriculum is generally easier to evaluate, supervise and improve than an open-ended educational chatbot. The findings therefore cannot automatically be generalised to every AI tutor currently marketed to schools or families.
The biggest unanswered question: does it scale?
The most important question for the broader AI bloom argument is not whether 165 students learned slightly more maths.
It is whether similar systems can retain their effectiveness when expanded by one or two orders of magnitude.
The research team appears to recognise this. Eedi and Google DeepMind have already launched a much larger follow-up trial involving 1,525 students across ten UK secondary schools. That study is designed partly to test whether the earlier findings survive at greater scale and over a longer period.[eedi.com]eedi.comstudy running with 1,525 students across 10 UK…Read more…
Scale introduces challenges that small pilots often avoid:
- More variation in student ability.
- More variation in school quality.
- Different teacher attitudes.
- Greater opportunities for misuse.
- Higher operational costs.
- Potential widening of achievement gaps if some students benefit more than others.
Educational history is full of interventions that looked impressive in pilots and less impressive in large deployments. The next generation of studies will therefore be more informative than the original trial about whether AI tutoring can become a broadly useful educational infrastructure.
Why this evidence still matters for the larger AI future
The strongest interpretation of the Eedi study is neither triumphalist nor dismissive.
It does not show that AI has solved education. It does not show that teachers are becoming obsolete. It does not prove that AI tutors will generate enormous learning gains across society.
What it does show is that a frontier AI model, deliberately adapted for pedagogy and embedded within human oversight, can participate in tutoring interactions without obvious educational collapse and may even improve certain learning outcomes.[arXiv]arxiv.orgAI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025…
For people interested in the larger possibility of AI-enabled human flourishing, that is the real signal. A future of abundant cognitive support would not arrive through a single dramatic breakthrough. It would emerge through many smaller demonstrations that machines can help people think, learn and reason more effectively in specific domains.
The 165-student Eedi trial is one of those demonstrations. Its value lies less in proving that AI tutoring already works at civilisation scale and more in showing that the idea is plausible enough to justify much larger, more demanding tests. The study is best understood as an encouraging piece of evidence, not a final verdict.
Amazon book picks
Further Reading
Books and field guides related to How strong is the Learn LM evidence?. Use these as the next step if you want deeper reading beyond the article.
Make It Stick
Strong match for measuring genuine learning transfer rather than short-term performance.
The AI Classroom
Directly relevant to supervised AI tutoring and classroom deployment.
Endnotes
1.
Source: arxiv.org
Link:https://arxiv.org/abs/2512.23633
Source snippet
AI tutoring can safely and effectively support students: An exploratory RCT in UK classroomsDecember 29, 2025...
Published: December 29, 2025
2.
Source: eedi.com
Link:https://www.eedi.com/news/just-launched—our-second-ai-tutor-rct
Source snippet
study running with 1,525 students across 10 UK...Read more...
3.
Source: arxiv.org
Title: arXiv Learn LM: Improving Gemini for Learning
Link:https://arxiv.org/abs/2412.16429
Source snippet
LearnLM: Improving Gemini for LearningDecember 21, 2024...
Published: December 21, 2024
4.
Source: scale.stanford.edu
Title: Randomized Controlled Trial
Link:https://scale.stanford.edu/ai/repository/impact-randomized-controlled-trial
Source snippet
START HERE for the most rigorous research on...
5.
Source: arxiv.org
Link:https://arxiv.org/html/2512.23633v1
Source snippet
AI tutoring can safely and effectively support studentsOur RCT aimed to evaluate LearnLM in a rigorous, real-world, in-classroom testbed...
6.
Source: arxiv.org
Link:https://arxiv.org/html/2602.19303v1
Source snippet
The Path to Conversational AI Tutors22 Feb 2026 — In one randomized [control]({{ 'control/' | relative_url }}) trial, researchers found that a “content-rich prompt engineer...
7.
Source: arxiv.org
Link:https://arxiv.org/pdf/2603.20088
Source snippet
Towards a Robust Evaluation Methodology for AI Systems...by J Edgell · 2026 · Cited by 1 — Ai tutoring can safely and effectively suppor...
8.
Source: eedi.com
Link:https://www.eedi.com/news/gold-standard-rct-shows-eedi-delivers-two-months-of-additional-maths-progress-2
Source snippet
the impact on learning, delivering four additional months of academic...
9.
Source: eedi.com
Link:https://www.eedi.com/news/new-uk-study-finds-students-using-eedi-gain-2-4-months-of-additional-maths-progress
Source snippet
New UK Study Finds Students Using Eedi Gain 2–4...7 Nov 2025 — Independent randomised controlled trial across 20 schools and 3,000 stude...
10.
Source: eedi.com
Link:https://www.eedi.com/research
11.
Source: eedi.com
Link:https://www.eedi.com/tag/press-release
Source snippet
Press ReleaseA large-scale independent randomised controlled trial (RCT) reveals that secondary school students using Eedi, an evidence-b...
12.
Source: scale.stanford.edu
Title: The Evidence Base on AI in K 12 Report
Link:https://scale.stanford.edu/sites/default/files/The%20Evidence%20Base%20on%20AI%20in%20K-12%20Report.pdf
Source snippet
Randomized Controlled Trial. (RCT). Quasi-Experimental Design. (QED). Peer-Reviewed.Read more...
13.
Source: storage.googleapis.com
Title: learn LM nov25
Link:https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_nov25.pdf
Source snippet
Each student and each...
14.
Source: linkedin.com
Title: structure beats access five genai pedagogical designs dr rod lane s3sve
Link:https://www.linkedin.com/pulse/structure-beats-access-five-genai-pedagogical-designs-dr-rod-lane-s3sve
Source snippet
Structure beats access: Five GenAI pedagogical designs...AI tutoring can safely and effectively support students: An exploratory RCT in...
15.
Source: socialscienceregistry.org
Link:https://www.socialscienceregistry.org/trials/18079
Source snippet
Testing the efficacy of AI tutoring in secondary mathematics13 Apr 2026 — An exploratory randomised controlled trial (RCT) conducted in 2...
Additional References
16.
Source: linkedin.com
Link:https://www.linkedin.com/posts/benkornell_breakthrough-aiedu-aiforeducation-activity-7394047991826870272-mP3s
Source snippet
AI collaboration boosts learning outcomes, says Eedi and...This is incredible: "Students who received LearnLM (supervised) tutoring wer...
17.
Source: linkedin.com
Link:https://www.linkedin.com/posts/simon-woodhead_edtech-aied-edm-activity-7393960569541652481-kR0c
Source snippet
Eedi and Google DeepMind: Human-Supervised LearnLM...Henotace AI transforms how students study for WAEC by providing personalized learni...
18.
Source: scirate.com
Link:https://scirate.com/?date=2026-01-09&page=162&range=56
Source snippet
Top arXiv papers... human tutors on each learning outcome we measured. In fact, students who received support from LearnLM were 5.5 perce...
19.
Source: edtechinnovationhub.com
Title: eedi and google deepmind begin second ai tutoring trial across 1525 uk students
Link:https://www.edtechinnovationhub.com/news/eedi-and-google-deepmind-begin-second-ai-tutoring-trial-across-1525-uk-students
Source snippet
EdTech Innovation HubEedi and DeepMind trial AI tutor with 1525 UK students5 May 2026 — The 12-week trial, running with students in years...
Published: May 2026
20.
Source: the74million.org
Title: ai tutors with a little human help offer reliable instruction study finds
Link:https://www.the74million.org/article/ai-tutors-with-a-little-human-help-offer-reliable-instruction-study-finds/
Source snippet
AI Tutors, With a Little Human Help, Offer 'Reliable'...3 Dec 2025 — New UK research finds 'extremely personalized' AI math tutors don't...
21.
Source: learning-engineering-virtual-institute.org
Link:https://learning-engineering-virtual-institute.org/ai-tutoring-can-safely-and-effectively-support-students-an-exploratory-rct-in-uk-classrooms/
Source snippet
One-to-one tutoring is widely considered the gold standard for personalized education, yet it remainsRead more...
22.
Source: future-ed.org
Title: research notes two emerging strategies for using ai in tutoring
Link:https://www.future-ed.org/research-notes-two-emerging-strategies-for-using-ai-in-tutoring/
Source snippet
Research Notes: Two Emerging Strategies for Using AI in...17 Feb 2026 — One study, conducted by researchers from Google and Eedi Labs, e...
23.
Source: eedi-labs-67qwnm2lxu4blqi5.webflow.io
Title: 165 students and 17 tutors, was a careful
Link:https://eedi-labs-67qwnm2lxu4blqi5.webflow.io/news/new-exploratory-research-from-eedi-and-google-deepmind-reveals-human-in-the-loop-ai-tutoring-outperforms-human-only-support
Source snippet
New Exploratory Research from Eedi and Google...The trial took place in the summer of 2025 across five UK secondary school classrooms, u...
24.
Source: aiedresearcher.org
Link:https://aiedresearcher.org/page/4/
Source snippet
AI Ed ResearcherWe integrated LearnLM — a generative AI model fine-tuned for pedagogy — into chat-based tutoring sessions on the Eedi mat...
25.
Source: substack.nomoremarking.com
Link:https://substack.nomoremarking.com/p/maybe-llm-tutors-might-be-able-to
Source snippet
LLM tutors might be able to work...11 Jan 2026 — The study took place in 5 UK secondary schools whose students were used to using the Eed...
Topic Tree



