We sell an AI teacher for $29 a month, so you should read this page knowing exactly what our interest is. That said, we would rather answer this question honestly than watch you find the answer the expensive way, in March, after a year you cannot get back. So here is the real version, including the parts that are good for ChatGPT and bad for us. You can absolutely teach your child with a free chatbot, and a lot of families are doing it right now. The question is not whether it works. It is what specifically breaks, when, and whether the thing that breaks is the thing you actually needed.
First, the part nobody selling you something will say: it is genuinely good
A general purpose chatbot in 2026 can explain long division six different ways, write a passage at exactly a fourth grade reading level about the exact dinosaur your child cares about, generate twenty practice problems, check them, diagram a sentence, and be endlessly patient about all of it at eleven at night. Ten years ago a family would have paid a tutor forty dollars an hour for a worse version of that.
So if you are using it, keep using it. And use it better. Here are the four things that make the biggest difference, and they cost nothing:
- Tell it not to give answers. Say it explicitly, at the top, every session: "Do not give my child the answer. Ask one question at a time. If she is stuck, give the smallest possible hint and let her try again." This single instruction changes the behavior more than anything else you can do, and the section below explains why it matters more than it sounds like it should.
- Tell it who your child is. Age, grade, what they already understand, what they got stuck on last time, what they love. A chatbot that knows your son is eight, is solid on addition, panics at word problems, and is obsessed with sharks teaches a genuinely different lesson than one that knows nothing.
- Give it the actual book. If you are working through a curriculum, paste in or upload the pages. "Teach page 43" beats "teach fractions," because your child's book has an order and a vocabulary and the internet's average fraction lesson does not match it.
- Make it ask before it explains. "Before you teach anything, ask him two questions to find out what he already knows." Most people never do this, and it is the difference between a lecture and a lesson.
Do those four things and you will get more out of a free chatbot than most families ever do. We would rather you know that than not.
Now the thing that breaks first, and it is not what you would guess
It is not accuracy. Modern models get elementary and middle school material right the overwhelming majority of the time, and when they slip, a parent usually catches it.
The thing that breaks is that it helps too much.
Think about what a chatbot is built to be. It is an assistant. Its whole personality, the thing it was shaped to do across every other task you have ever given it, is to be maximally useful right now: answer the question, solve the problem, remove the friction. That is exactly what you want when you are writing an email at midnight.
It is exactly wrong for a child who is thirty seconds away from figuring something out.
Learning happens in that thirty seconds. The technical name for it is productive struggle, and every teacher knows it by feel: the moment where a kid is stuck, uncomfortable, and on the edge of getting it, and the right move is to say almost nothing. A good teacher sits on their hands. An assistant, by design and by instinct, jumps in.
So your child asks a question, and the chatbot gives a beautiful, clear, correct explanation. Your child says "oh, okay." The session felt great. Everyone was pleasant. And nothing was learned, because the hard part got done by the machine, and the hard part was the part that was supposed to happen inside your child's head.
This is the failure that compounds quietly, because there is no error message. A wrong answer you would catch. A child who is being smoothly, cheerfully helped past every difficulty for eight months looks like a child having a great school year, right up until somebody asks them to do it alone.
You do not have to take our word for it. Somebody with no stake in this measured it
In 2026 the Allen Institute for AI, a nonprofit research institute, released a benchmark called TutorMoments. It exists to answer a question the field had mostly avoided: not whether an AI can solve the math, but whether it knows when to step in and when to hold back.
Here is how they built it, because the method is the reason to trust the result. They took 462 de-identified transcripts of real one on one math tutoring with American students in grades 2 through 7, from a high dosage tutoring program serving mostly Title I schools. Twenty-seven experienced US math teachers went through those sessions and marked more than 1,500 key moments, the specific decision points where a tutor had to choose between giving support and pushing the student to do more of the thinking. Then they paused the transcript at those moments, handed the tutoring over to a language model for five turns with a simulated student, and scored whether the model did what the teachers said the moment called for.
The headline finding: told only to tutor well, the models over-help. They scaffold heavily and they rarely push for rigor. In the published results one model scored 0.738 on appropriate scaffolding and 0.181 on appropriate rigor under a plain prompt. That is a machine that is very good at helping and nearly incapable of holding back, which is precisely the failure described above, measured.
And the second finding matters even more. When the researchers changed nothing but the prompt, spelling out the scaffolding versus rigor trade-off explicitly, every score went up, in some cases dramatically. The same underlying model behaves like a different tutor depending on what it was told a tutor is for.
Read those two findings together and you have the whole argument of this page. The default behavior of a general assistant is the wrong behavior for a teacher, and the instructions are the difference. A chatbot with no instructions is not a bad teacher because it is a bad model. It is a bad teacher because nobody told it to be one, and its factory setting is "be helpful," which in a classroom means "do the child's thinking for them."
Three honest caveats, because a page that only quotes the convenient half of a study is not worth reading. Ai2 has never heard of us. We have not been measured by TutorMoments, we have no score on it, and nothing here should be read as a ranking, a result or an endorsement of anything we sell. Second, the researchers are careful about their own limits: the dataset is US elementary and middle school math, annotated by one pool of teachers, and they say plainly that an automated evaluation "can't stand in for studies with real students." Third, and this is the finding least convenient for everybody: the human tutors in that same data did not run away with it either. On rigor moments they scored low too, and the researchers deliberately treat human performance as a naturalistic reference rather than a ceiling, partly because the moments were chosen as places tutoring could have gone better. Nobody comes out of that paper looking like a genius. What comes out of it is a clear, independently documented statement of what the hard part is.
And it is not one study
A second benchmark landed in the same stretch, built completely differently, and it points the same direction. TutorBench was built by Scale AI, a commercial company rather than a nonprofit, so weigh it accordingly. It uses 1,500 student and tutor conversations across six high school and AP STEM subjects, more than half of them including photographs of actual student work, scored against a rubric of more than fifteen thousand human-written criteria.
Two things in it are worth your time. First, the rubric assigns its heaviest penalty to behaviors like "giving away a final answer when only a hint is requested," which is the same failure described above, written into a scoring system by people who had to decide what bad tutoring is. Second, the result: across fifteen frontier models, the best score reported was 55.65 percent, and the authors' summary is that no model has yet mastered the complexities of tutoring.
That number deserves to be read carefully rather than as a scare statistic, because a rubric score is not a grade on whether your child learns anything. What it fairly supports is a modest claim, and it is the claim of this whole page: tutoring is a separate skill from knowing the material, the models are not automatically good at it, and how a model is set up to behave is a real variable rather than a marketing word. The same honesty rules apply here as above. Scale AI has never heard of us either, we have not been evaluated by TutorBench, and we have no score on it.
The other four things a school year needs that a chat window does not have
Over-helping is the first failure. These are the four that show up between month two and month eight.
1. It does not know where your child actually is. A chat window starts at zero every time. It teaches whatever you asked about, at whatever level it guesses. It has no idea that your fifth grader is reading at a seventh grade level and doing third grade math, which is an extremely common shape for a real child and the single most important fact about teaching them. If nobody establishes where the child is, every lesson afterward is aimed at an average that does not exist.
2. It forgets. Memory features have improved and they are still not a school record. Ask yourself what a good teacher knows about your child in April that they did not know in September: that he shuts down when a page has too many problems on it, that she needs to say it out loud before she can write it, that the word "borrowing" specifically is the thing that never clicked. That knowledge accumulates in a teacher over a year and it is the reason a September teacher and an April teacher are not the same person. A conversation that resets does not accumulate anything.
3. Nothing writes it down. This is the one families discover late and painfully. Depending on your state, you may need an attendance record, an annual assessment on file, a portfolio, a log of instruction, or a transcript that a college will read. Even in a state that asks for nothing, you will eventually need evidence for somebody: a district placement conversation, a scholarship application, a grandparent, or yourself in a bad week. Chat history is not a record. Nobody is going to accept a folder of conversation links, and nobody wants to reconstruct 180 days from them.
4. There is no year. A school year is not a pile of good lessons. It is a sequence, paced across roughly 180 days, where things come in an order for a reason and there is a plan for what happens if week nine goes badly. A chat window has no calendar, no scope, no sequence and no memory of what you covered in October. The pacing job stays entirely on you, and pacing is most of the job that exhausts homeschooling parents.
And the quiet fifth one: there is no gate. Nothing in a chat stops a child from nodding along. They can say "yeah I get it" and move on, forever, and the machine will cheerfully agree, because agreeing is what it does.
So what would you actually have to build
Look at that list and notice something: none of it is about the model being smarter. Every single item is about what surrounds the model. A teacher that holds back on purpose. A real placement conversation before the first lesson. A profile of your specific child that survives the year. A record that writes itself. A calendar. A gate that will not let a child move on before they can actually do it.
That is what we built, and that is the entire honest pitch. We are not claiming a better brain. We are claiming the scaffolding around it. The teacher is told, in detail, what a teacher is for and when to shut up. It meets your child in a free placement conversation first, one that feels like a friendly chat rather than a test, and tells you where they actually are in each subject instead of where the grade label says. It keeps a running profile of how your child learns and carries it forward. It teaches from your curriculum if you have one, paced across your year. Every session writes itself down: date, subject, book and unit, minutes, what was worked on, what got solid. Review days check whether things actually stuck. And the record it produces prints, because a state form does not accept a screenshot.
We also publish what changed each night on our improvements page, in plain language, so you can see the thing getting better instead of being told it does.
The elephant, said out loud
Yes, we are describing a product built on the same kind of model you can talk to for free. That is a fair thing to be suspicious of, so here is the straight version.
What you get free is the engine. What $29 a month buys is everything that turns an engine into a school year: the instructions that make it teach instead of answer, the placement, the memory of your child, the pacing, the mastery gate, the printable record, and the fact that somebody is improving it every night whether or not you have the time to. You could assemble a version of that yourself with enough prompting discipline, and a handful of technically minded parents genuinely do. Most people, running a house with two or three kids in it, will not do that consistently for 180 days. That is not a knock on you. It is the reason the product exists.
And here is the part we will hold ourselves to: if we are ever measured on something like TutorMoments, we will publish the number whatever it says. This page would be worthless if we only ever cited research that made us look good.
The test you can run tonight, free, with no card
Do not take our word or anyone else's. Run this, it takes twenty minutes.
- Pick something your child is genuinely stuck on. Not something they know. The stuck thing.
- Open a free chatbot and ask it to tutor them. Do not give it special instructions the first time. Watch what happens in the first four exchanges: count how many times it explains, and how many times it makes your child do the thinking.
- Now start over and give it the instruction from the top of this page: no answers, one question at a time, smallest possible hint. Watch the difference. This is the TutorMoments finding happening on your kitchen table.
- Then ask the question nobody's chat window can answer: "Where is my child, right now, in reading and math, and what should we do Monday?"
That last one is the whole thing. If you can get a real answer to it, from anywhere, you are ahead of most families. That is the question our free placement assessment exists to answer, and it costs nothing and asks for no card, because it is also the question we think you deserve an answer to whether or not you ever pay us a dollar.
Related reading: how to figure out the way your child actually learns, whether you are qualified to homeschool (you almost certainly are), and how to choose a curriculum without spending a fortune. Or see the homeschool requirements for your state.
Answer the one question a chat window can't
Where is your child actually standing right now, in each of the five core subjects? The free placement assessment finds out. It feels like a friendly chat, never a test, it takes about fifteen minutes a subject and you can do one a day, and you get a real answer whether or not you ever sign up. No card.
Start with the free assessment