PolyzPolyz
All articles

Do AI Tutors Actually Work? What the Research Says

Every AI tutor cites Bloom's two-sigma study, and most cite it wrong. What the tutoring research really found, and how to spot a real AI tutor.

Polyz Team8 min read

Almost every AI tutor on the market cites the same study. You've seen the claim even if you didn't catch the source: one-on-one tutoring makes the average student perform better than 98% of a normal classroom. Two standard deviations. The "two sigma problem." It gets used as a syllogism — tutoring produces two sigma, this app is a tutor, therefore buy this app.

The study is real. The syllogism is not. And the genuinely interesting part is that when you go read the actual literature, the case for AI tutoring is still good — just for different reasons, with different numbers, and with a specific set of conditions that most products fail to meet.

Here's what the research says, including the parts that don't flatter the industry.

The two-sigma claim is the weakest evidence, not the strongest

Benjamin Bloom published "The 2 Sigma Problem" in 1984. His graduate students ran experiments comparing conventional classes, mastery learning, and one-to-one tutoring combined with mastery learning. The tutored group scored about two standard deviations above the conventional group. Bloom framed it as a challenge: find a scalable method that gets the same result.

The problem is what happened next, which is essentially nothing. That effect size has never been reproduced at anything close to that magnitude. The original experiments ran on small groups of schoolchildren, over units measured in weeks, on material the tests were written to cover closely. Tutoring works. It does not reliably work that well, and a forty-year-old result that resisted replication is a strange thing to build a product claim on.

If a company's homepage leads with two sigma, that tells you something about their marketing department and nothing about their software.

The number that actually matters

The more useful study is Kurt VanLehn's 2011 meta-analysis in Educational Psychologist, which compared human tutors, intelligent tutoring systems, and simpler computer-based tutors against no tutoring at all.

The received wisdom going in was a tidy hierarchy: answer-based computer tutors around d = 0.3, intelligent tutoring systems around d = 1.0, human tutors around d = 2.0. What VanLehn found instead:

  • Human tutoring: d = 0.79
  • Intelligent tutoring systems: d = 0.76
  • Answer-based computer tutoring: d = 0.3

Two things fall out of this, and they point in opposite directions.

The first is deflationary: human tutors are worth about 0.79, not 2.0. The gold standard is less golden than the folklore.

The second is the actually remarkable finding. Software scored 0.76 against a human's 0.79. The gap between a well-built tutoring system and a live human being was, statistically, close to nothing — and this was in 2011, using hand-authored systems built years before large language models existed. Those systems were narrow. They covered one topic, in one course, with rules written by hand for every step a student might take. They were enormously expensive to build, which is exactly why they never scaled beyond algebra and physics.

The interesting question was never "can software tutor." It was answered fifteen years ago. The question is whether language models make that capability cheap and general enough to point at anything.

What the LLM trials show

Two results from the last two years are worth knowing about.

Harvard, first-year physics. Gregory Kestin and Kelly Miller ran a randomized study with 194 students in a physics course for life-sciences majors. One group attended the usual class — and this matters, because Kestin's usual class is active learning, where students work problems in groups with live feedback, which is already the format that beats lecturing in the research. The other group worked with a custom AI tutor called PS2 Pal in their dorms.

The AI group learned roughly twice as much, in less time.

The design detail is the whole story. PS2 Pal was not a chatbot with a friendly name. It was constrained: keep replies to a few sentences to avoid overloading working memory, give away one step at a time, never hand over a complete solution. They built the pedagogy into the constraints. A raw model asked the same questions does the opposite — it produces a clear, complete, beautifully-organized answer, which is the single most reliable way to prevent someone from learning.

Edo State, Nigeria. The World Bank ran a randomized controlled trial with around 800 senior secondary students across nine public schools in Benin City. Six weeks, twelve 90-minute after-school sessions, teachers guiding students through GPT-4 via Microsoft Copilot on English grammar and writing.

The treatment group significantly outperformed the control — including on end-of-year exams covering material the intervention never touched, which is a much harder result to fake than a test written to match the lesson. Cost was about $48 per student.

You will see this cited as "two years of learning in six weeks" and as "2.23 standard deviations." Be careful with both. The 2.23 figure is an extrapolation of six weeks of gains out to a full academic year, not something anyone measured. Extrapolating a short intervention linearly is exactly where novelty effects hide. The measured six-week result is genuinely strong. The annualized headline is a projection, and the researchers say so.

The caveats, stated plainly

If you only read the wins, you'll build the wrong expectations.

  • The trials are short. Six weeks, one semester, one unit. Nobody has good data on what an AI tutor does for you over two years, which is the timeframe most adults actually care about.
  • Almost everyone studied is a student. They're enrolled in a course, with a syllabus, a deadline, and an exam. That external structure is doing work that nobody is measuring separately. A self-directed adult learning statistics on weeknights has none of it.
  • Novelty inflates early results. New thing, more attention, better engagement, better scores. Some fraction of every early-stage edtech result is this, and it decays.
  • The successes were engineered. PS2 Pal beat an active-learning classroom because someone deliberately stopped it from behaving like a normal assistant. This is not a property of language models. It is a property of that build.
  • Publication bias is real. Trials that find nothing are harder to publish and much harder to press-release.

None of this makes the evidence weak. It makes it specific. The claim the research supports is not "AI tutors work." It's "systems built around a particular set of teaching behaviors work, and language models make those systems cheap to build for any subject."

The four behaviors that separate teaching from answering

Strip the studies down and the same mechanisms keep appearing. These are worth knowing as a buyer, because they're what you should be checking for.

1. It measures before it teaches. Mastery learning — the other half of Bloom's experiment, the half nobody quotes — means you don't advance until you've demonstrated the current thing. That requires knowing where you're starting. A tutor that begins teaching before it has any idea what you already know is guessing, and it will spend your first three sessions on material you didn't need.

2. It withholds the answer. The retrieval-practice literature is about as settled as education research gets: struggling to produce an answer builds durable memory, and reading a correct answer produces a warm feeling of understanding that evaporates within days. This is the single most common failure of general-purpose chatbots, and it isn't a bug in the model. Being maximally helpful right now means being maximally unhelpful to your memory next week. PS2 Pal's "one step at a time, never the full solution" rule exists precisely to defeat the model's default.

3. It grades your work, not just its own. There's a large difference between a system that answers your questions and one that evaluates your attempts. Only the second can find out what you don't know — including the things you don't know you don't know, which are the ones that actually block you.

4. It remembers and adapts. Everything above is worthless if it resets each session. The gap you had on Tuesday should shape Thursday's lesson. This is the ordinary meaning of "adaptive," as opposed to the marketing meaning, which is usually "the chat has scrollback."

How to evaluate any AI tutor in ten minutes

Ignore the landing page. Run these four checks:

  • Ask it to teach you something, then ask a question you know the answer to. Does it explain and hand you the answer, or does it ask you something back? If you get a clean, complete, well-formatted answer every time, you have a search engine with good manners.
  • Deliberately get something wrong. A tutor should identify which misconception you have, not just mark it incorrect and restate the right answer.
  • Come back the next day. Does it know what you struggled with, or are you starting over? If you have to re-explain your level, nothing is being tracked.
  • Ask what your syllabus is. A real course knows what comes after this. A chat knows what you just typed.

Most tools fail the first check, and it's the cheapest one to run.

Where this leaves you

The honest summary: tutoring produces a large, well-replicated effect somewhere around d = 0.8 — not two sigma. Software has matched human tutors on that measure since at least 2011, back when each system had to be hand-built for a single topic. Language models remove that constraint, and the early randomized trials on LLM-based tutoring are encouraging enough to take seriously and short enough that you shouldn't extrapolate them.

What decides whether a given product works is not the model behind it. Every serious tool is running one of a handful of frontier models. It's whether the thing was built to teach — to measure first, withhold answers, grade attempts, and remember — or built to be helpful, which is a different and much easier goal that happens to undercut learning.

That distinction is why we built Polyz as an AI tutor rather than another chat window: it runs a diagnostic before it teaches you anything, builds an actual syllabus for your goal, grades your practice, and rewrites what comes next based on what you got wrong. If you want the mechanics of how that works in practice, the comparison of the four ways to learn something with AI covers where each one breaks down.

And if you want to try the approach before paying for anything, the fastest version is free: read how to use ChatGPT as a tutor, which includes the prompt that gets a general chatbot to behave like PS2 Pal — and an honest account of the four places it still falls over.


Sources: Bloom, B. S. (1984), "The 2 Sigma Problem," Educational Researcher. VanLehn, K. (2011), "The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems," Educational Psychologist 46(4). Kestin, G. & Miller, K. et al. (2024), AI tutoring vs. active learning in Harvard PS2. World Bank Education Global Practice (2024), Edo State, Nigeria generative-AI tutoring RCT.

Try Polyz free for 7 days

Every feature, all three Claude models, a generous monthly quota. Card required to start. Cancel before day 7 and you owe nothing.