Within Education

When should an AI tutor refuse the answer?

AI tutors can help learners practise, but poorly designed systems can let students outsource the very effort that builds skill.

On this page

  • The crutch problem in homework and practice
  • Hints, retrieval and Socratic questioning as safeguards
  • How schools can test whether learning transfers
Preview for When should an AI tutor refuse the answer?

Introduction

An AI tutor that instantly solves every problem can look impressive, but it may undermine the very thing education is trying to build. Learning is not just producing the correct answer. It is acquiring skills that still work when the tutor disappears.

Tutor Guardrails illustration 1 This is why one of the most important design questions in AI education is surprisingly simple: when should the system refuse to answer? The strongest case for AI-driven cognitive empowerment is not that students can outsource thinking more efficiently. It is that billions of people could gain access to guidance that helps them think better themselves. That requires guardrails which keep learners mentally active rather than turning AI into a homework-completion machine. Research increasingly suggests that these design choices matter. Systems that provide unrestricted solutions can improve assignment performance while weakening independent mastery, whereas tutors built around hints, questioning and staged support appear far more compatible with genuine learning.[PMC]pmc.ncbi.nlm.nih.govGenerative AI without guardrails can harm learning - PMC - NIHby H Bastani · 2025 · Cited by 198 — Our research examines the impact of…

The crutch problem in homework and practice

The central danger is not that students will occasionally cheat. It is that AI can quietly replace the cognitive effort that creates understanding.

A large field experiment in high-school mathematics examined what happened when students used GPT-4 during practice sessions. Students often performed better while the AI was available, but many learned less effectively when later required to work independently. The researchers described a recurring pattern: learners used the model as a “crutch”, obtaining solutions instead of building the skills needed to solve similar problems themselves. pnas.org PubMed This problem is easy to misunderstand. A student can appear productive while learning very little. If an AI writes the equation[khanmigo.ai]khanmigo.aiYou can view your child's chats, get alerts for flagged content, and feel good…Read more…, selects the method, performs the calculation and explains the result, the learner may experience a feeling of comprehension without having done the retrieval, reasoning and error-correction that durable learning requires.

The distinction matters for the broader AI bloom vision. A future of abundant intelligence is not achieved merely because more answers are available. Search engines already made answers abundant. The stronger promise is that advanced AI could help people acquire knowledge, judgement and problem-solving ability at far larger scale. If educational AI encourages cognitive dependency instead, then some of the apparent productivity gains may come at the cost of reduced human capability.

Several researchers have therefore argued that educational AI should be evaluated differently from ordinary chatbots. The question is not simply whether the answer is correct or safe. It is whether the interaction strengthens the learner’s future ability to think independently.[arXiv]arxiv.orgarXiv Safe Tutors: Benchmarking Pedagogical Safety in AI Tutoring SystemsSafeTutors: Benchmarking Pedagogical Safety in AI Tutoring SystemsMarch 18, 2026…Published: March 18, 2026

When should an AI tutor refuse the answer?

The most important guardrail is often selective refusal.

In traditional tutoring, a good teacher does not automatically reveal the solution the moment a student struggles. They try to identify where understanding breaks down and then provide the minimum help needed to move learning forward.

Many AI tutoring projects are attempting to replicate this principle. Khan Academy’s Khanmigo explicitly promotes a Socratic approach, guiding students through questions rather than supplying direct answers. Its public descriptions repeatedly emphasise critical thinking, step-by-step reasoning and resisting requests for immediate solutions.[edutopia]edutopia.orgai tutors work guardrailsAI Tutors Can Work—With the Right GuardrailsMar 27, 2025 — As Khan Academy shows in its demo, both have been designed not to give student… The key idea is not refusal for its own sake. A tutor that endlessly withholds information can become frustrating and useless. Instead, effective guardrails try to answer a different question:

What is the smallest intervention that helps the learner make progress themselves?

That can mean:

  • Asking what the student already knows.[linkedin.com]linkedin.comWhen a student asks "what's the answer?" — Khanmigo responds with "What do you already know about this problem?" and guides…Read more…
  • Requesting the next step rather than solving the whole problem.
  • Checking for misconceptions before providing guidance.
  • Revealing one hint at a time.
  • Requiring an attempted answer before offering stronger support.
  • Explaining why an approach works rather than merely displaying it.

The mathematics study from the University of Pennsylvania and collaborators found that these kinds of safeguards substantially reduced the negative learning effects associated with unrestricted GPT-4 use. Students using a guarded tutor retained more independent capability than those using a system that freely delivered answers.[SSRN]papers.ssrn.comAI Without Guardrails Can Harm Learningby H Bastani · 2024 · Cited by 344 — Without guardrails, students attempt to use GPT-4 as a “crutc…

Hints, retrieval and Socratic questioning as safeguards

Learning science has long shown that mental effort matters. People remember information better when they must retrieve it from memory, explain it in their own words or apply it to a new problem.

That insight translates directly into AI tutor design.

Retrieval before explanation

One simple safeguard is forcing recall before assistance.

Instead of saying, “Here is the formula”, the tutor might ask:

  • What formula did your class use for problems like this?
  • Can you remember any part of it?
  • Which quantities are you trying to relate?

This creates retrieval practice, a well-established mechanism for strengthening memory. The learner must actively reconstruct knowledge rather than passively receiving it.

Layered hints

Another guardrail is staged disclosure.

Rather than exposing the entire solution immediately, the tutor provides progressively stronger support:

  1. A conceptual hint.
  2. A procedural hint.
  3. A worked intermediate step.
  4. The full solution only if necessary.

This preserves productive struggle while preventing frustration from becoming overwhelming.

Tutor Guardrails illustration 2

Socratic dialogue

Perhaps the most discussed approach is Socratic questioning.

Instead of acting like an answer engine, the AI behaves more like a tutor who continually asks:

  • Why do you think that?
  • What assumption are you making?
  • How could you check that result?
  • What would happen if the variable changed?

Recent work on critical-thinking assistants and Socratic AI systems argues that questioning can help students inspect their assumptions and develop reasoning skills rather than merely collecting conclusions.[arXiv]arxiv.orgarXiv Safe Tutors: Benchmarking Pedagogical Safety in AI Tutoring SystemsSafeTutors: Benchmarking Pedagogical Safety in AI Tutoring SystemsMarch 18, 2026…Published: March 18, 2026[2edtechbooks.org]edtechbooks.orgce of Learning. Andy Van Schaack & Roman Sarlo.Read more…

The broader significance is that these methods shift AI from cognitive offloading towards cognitive scaffolding. The system supports thinking instead of replacing it.

Why guardrails are harder than they look

Building these safeguards is not as simple as instructing a model: “Don’t give answers.”

Students are inventive.

Research on AI teaching assistants in programming courses found that many learners actively seek ways around guardrails, especially when deadlines approach or when they feel stuck. In one study, students were given the option to disable scaffolding and view unrestricted solutions. Roughly half used the feature at least once, and lower-performing students were especially likely to rely on it. Time pressure, lack of self-regulation and the desire for immediate completion were major drivers.[arXiv]arxiv.orgarXiv Safe Tutors: Benchmarking Pedagogical Safety in AI Tutoring SystemsSafeTutors: Benchmarking Pedagogical Safety in AI Tutoring SystemsMarch 18, 2026…Published: March 18, 2026

This reveals a deeper tension.

The educational goal is learning. The student’s immediate goal is often finishing the assignment.

When those goals diverge, many learners will choose short-term completion.

As a result, successful guardrails cannot rely entirely on student goodwill. They often need structural support from course design itself. A system that encourages reasoning may still fail if every assessment rewards answer production alone.

Researchers developing tutoring benchmarks have increasingly argued that pedagogical safety differs from ordinary AI safety. The main risk is not offensive content or factual errors. It is gradual learning erosion: answer over-disclosure, misconception reinforcement and excessive dependence on the model. Multi-turn conversations can make these failures worse because persistent interaction gives learners more opportunities to extract solutions indirectly.[arXiv]arxiv.orgarXiv Safe Tutors: Benchmarking Pedagogical Safety in AI Tutoring SystemsSafeTutors: Benchmarking Pedagogical Safety in AI Tutoring SystemsMarch 18, 2026…Published: March 18, 2026

Tutor Guardrails illustration 3

How schools can test whether learning transfers

The simplest way to evaluate an AI tutor is to ask whether students complete more work.

That is also one of the weakest tests.

A better question is whether learning transfers beyond the AI-assisted environment.

Schools and universities increasingly need assessment methods that distinguish genuine understanding from successful AI use.

Useful indicators include:

Performance without AI access. If scores collapse when assistance is removed, the tutor may be supporting output rather than learning. This was one of the key findings from the mathematics field experiment.[pnas.org]pnas.orgGenerative AI without guardrails can harm learningby H Bastani · 2025 · Cited by 151 — Without guardrails, students attempt to use GPT-4…

Novel problem solving. Students should be able to apply concepts to unfamiliar questions rather than merely reproducing solutions they have already seen.

Explanation quality. Learners can be asked to explain why an answer works, not merely state the answer itself.

Delayed retention. Knowledge measured weeks later often provides a clearer picture of learning than immediate homework performance.

Oral reasoning and discussion. Teachers can probe understanding through conversation, making it harder to hide behind AI-generated work.

The more education systems reward transferable reasoning, the more incentive AI tutors will have to cultivate it.

The larger question for AI-enabled education

The guardrail debate points to a larger issue inside the AI bloom vision.

If advanced AI becomes vastly more capable than today’s systems, the temptation to outsource thinking will grow. Future models may be able to solve scientific problems, write persuasive essays, generate software and perform complex analysis with very little human effort.

That creates two possible futures for education.

In one, people increasingly depend on machine intelligence while their own skills stagnate. Productivity rises, but human cognitive agency gradually weakens.

In the other, AI functions more like a universal intellectual coach: helping people understand faster, practise more effectively, overcome educational barriers and participate in increasingly sophisticated forms of work and discovery.

The difference may depend less on raw model capability than on design choices. A tutor that instantly completes homework demonstrates intelligence abundance. A tutor that helps a learner become more capable demonstrates cognitive empowerment.

For education to contribute to long-term human flourishing rather than mere automation, AI systems will need to preserve a difficult balance: helpful enough to unlock learning, but resistant enough to keep the learner doing the work that learning requires. Research increasingly suggests that this balance is not an optional feature. It is the core mechanism that determines whether AI tutoring expands human capability or quietly replaces it.[2chibe.upenn.edu]chibe.upenn.eduGenerative AI without guardrails can harm learningThis study tested generative AI tutors, showing that design guardrails, or prompts that…

Amazon book picks

Further Reading

Books and field guides related to When should an AI tutor refuse the answer?. Use these as the next step if you want deeper reading beyond the article.

BookCover for Make It Stick

Make It Stick

By Peter C. Brown, Henry L. Roediger III et al.

Directly supports the need for retrieval, effort and transfer rather than instant answers.

eBay marketplace picks

Marketplace Samples

Example marketplace items related to this page. Use the search link to explore similar finds on eBay.

UsingUSA

Endnotes

1. Source: pnas.org
Link:https://www.pnas.org/doi/10.1073/pnas.2422633122

Source snippet

Generative AI without guardrails can harm learningby H Bastani · 2025 · Cited by 151 — Without guardrails, students attempt to use GPT-4...

2. Source: pmc.ncbi.nlm.nih.gov
Link:https://pmc.ncbi.nlm.nih.gov/articles/PMC12232635/

Source snippet

Generative AI without guardrails can harm learning - PMC - NIHby H Bastani · 2025 · Cited by 198 — Our research examines the impact of...

3. Source: chibe.upenn.edu
Link:https://chibe.upenn.edu/publications/generative-ai-without-guardrails-can-harm-learning-evidence-from-high-school-mathematics/

Source snippet

Generative AI without guardrails can harm learningThis study tested generative AI tutors, showing that design guardrails, or prompts that...

4. Source: arxiv.org
Title: arXiv Safe Tutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
Link:https://arxiv.org/abs/2603.17373

Source snippet

SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring SystemsMarch 18, 2026...

Published: March 18, 2026

5. Source: arxiv.org
Link:https://arxiv.org/html/2605.04816v2

Source snippet

Building AI Companions that Prioritise Learning over...15 May 2026 — Critical Thinking Assistants use a form of Socratic questionin...

Published: May 2026

6. Source: khanmigo.ai
Link:https://www.khanmigo.ai/

Source snippet

Meet Khanmigo: Khan Academy's AI-powered teaching...Khanmigo challenges you to think critically and solve problems without giving you di...

7. Source: edutopia.org
Title: ai tutors work guardrails
Link:https://www.edutopia.org/article/ai-tutors-work-guardrails/

Source snippet

AI Tutors Can Work—With the Right GuardrailsMar 27, 2025 — As Khan Academy shows in its demo, both have been designed not to give student...

8. Source: khanmigo.ai
Link:https://www.khanmigo.ai/learners

Source snippet

Khanmigo for learners: Always-available tutor, powered by AIDiscover a new way to learn, powered by AI. Khanmigo is your always-available...

9. Source: papers.ssrn.com
Link:https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4895486

Source snippet

AI Without Guardrails Can Harm Learningby H Bastani · 2024 · Cited by 344 — Without guardrails, students attempt to use GPT-4 as a “crutc...

10. Source: edtechbooks.org
Link:https://edtechbooks.org/promptbook/from-oracle-to-socratic-partner

Source snippet

ce of Learning. Andy Van Schaack & Roman Sarlo.Read more...

11. Source: arxiv.org
Link:https://arxiv.org/abs/2504.11146

Source snippet

Exploring Student Behaviors and Motivations using AI TAs with Optional GuardrailsApril 15, 2025...

Published: April 15, 2025

12. Source: knowledge.wharton.upenn.edu
Title: without guardrails generative ai can harm education
Link:https://knowledge.wharton.upenn.edu/article/without-guardrails-generative-ai-can-harm-education/

Source snippet

Guardrails, Generative AI Can Harm EducationAug 27, 2024 — Students who rely on generative AI to help them learn may be missing out on ba...

13. Source: pnas.org
Title: Learning is critical to long-term productivity
Link:https://www.pnas.org/doi/abs/10.1073/pnas.2422633122?gad_campaignid=21506599862&gad_source=1&gbraid=0AAAAAqiPmCx-8XoPPf3prdOo4pSzOXIvP&gclid=CjwKCAjwyMnNBhBNEiwA-Kcgu_cdG0PO6KphSETT1AQt7CwWIr4Q76HDRebsxVrHKOqzg-LxEEkLaxoCRYUQAvD_BwE

Source snippet

Generative AI without guardrails can harm learningby H Bastani · 2025 · Cited by 151 — A key question is how generative AI affects learni...

14. Source: khanmigo.ai
Link:https://www.khanmigo.ai/parents

Source snippet

You can view your child's chats, get alerts for flagged content, and feel good...Read more...

15. Source: pubmed.ncbi.nlm.nih.gov
Link:https://pubmed.ncbi.nlm.nih.gov/40560616/

Source snippet

AI without guardrails can harm learningby H Bastani · 2025 · Cited by 147 — Without guardrails, students attempt to use GPT-4 as a "crutc...

Additional References

16. Source: linkedin.com
Link:https://www.linkedin.com/posts/hamsa-bastani-4a346955_generative-ai-without-guardrails-can-harm-activity-7343667696540033025-a1Ms

Source snippet

How generative AI affects learning in mathOur research examines the impact of generative AI, specifically GPT-4, on student learning in m...

17. Source: facebook.com
Link:https://www.facebook.com/khanacademy/posts/did-you-hear-the-news-openais-newest-model-can-reason-across-audio-vision-and-te/857641319740336/

Source snippet

Khan AcademyAdjacent was an answer despite ai was asked to do not give an answer. Opposite was the only remaining option out of three, th...

18. Source: facebook.com
Link:https://www.facebook.com/groups/703007927897194/posts/1283818296482818/

Source snippet

Using AI Tools to Enhance Student Writing**What Teachers Should Know About ChatGPT’s New Study Mode Feature** This post describes how a n...

19. Source: medium.com
Link:https://medium.com/%40adnanmasood/the-quiet-math-of-edtech-can-ai-tutors-really-teach-737195005abf

20. Source: linkedin.com
Link:https://www.linkedin.com/posts/smart-staff-room_most-ai-tools-give-students-the-answer-activity-7442506264066084864-CXwK

Source snippet

When a student asks "what's the answer?" — Khanmigo responds with "What do you already know about this problem?" and guides...Read more...

21. Source: cbsnews.com
Link:https://www.cbsnews.com/news/khanmigo-ai-powered-tutor-teaching-assistant-tested-at-schools-60-minutes-transcript/

Source snippet

AI-powered tutor tested as a way to help educators and...Dec 8, 2024 — It's an online tutor powered by artificial intelligence designed...

22. Source: blog.khanacademy.org
Title: how khan academy is building a better ai tutor our most recent learnings
Link:https://blog.khanacademy.org/how-khan-academy-is-building-a-better-ai-tutor-our-most-recent-learnings/

Source snippet

Khan Academy Is Building a Better AI Tutor: Our Most...May 1, 2026 — Khan Academy shares how it improved its AI tutor Khanmigo with fast...

Published: May 1, 2026

23. Source: freethink.com
Title: Sal Khan wants to give every student on Earth a personal AI tutor
Link:https://www.freethink.com/consumer-tech/khanmigo-ai-tutor

Source snippet

January 25, 2025 — This Socratic approach extends to Khanmigo Writing Coach, a generative AI-based tool that Khan Academy developed speci...

Published: January 25, 2025

24. Source: hamsabastani.github.io
Link:https://hamsabastani.github.io/education_llm.pdf

Source snippet

4, on student learning in math education. Through a large-scale field experiment in a high.Read more...

25. Source: rickhess99.medium.com
Title: can an ai powered tutor produce meaningful results b67d7376cb51
Link:https://rickhess99.medium.com/can-an-ai-powered-tutor-produce-meaningful-results-b67d7376cb51

Source snippet

an AI-Powered Tutor Produce Meaningful Results?If we said, “You are a Socratic tutor. I am a student. Don't give me answers to my questio...

Topic Tree

Follow this branch

Parent topic

Education Can Everyone Have a World Class Tutor?

Related pages 3

More on this topic 3