AI lesson tools aren't as accurate as you think. We checked 601 AI-made teaching resources to find out why.
AI writes like an expert. The sentences are fluent, the tone is confident, the layout looks finished. That is exactly what makes its mistakes easy to miss. Large language models are known to produce text that sounds right and isn't: the UK's Department for Education warns that generative AI "can produce nonsensical, inaccurate or false information presented as fact, known as hallucination".[1]
We wanted to know what that looks like in real lessons, so we started with our own. We read every one of the 601 resources in LessonHQ's official library, all written by earlier versions of our own AI generator, and logged every error. This page shows what we found, why it happens, what to check in any AI-made resource, and what we changed.
Whose mistakes these are
- Every number on this page comes from our own library: 601 resources (405 UK, 196 US) made by earlier versions of our own generator. We are publishing our mistakes, not anyone else's.
- We did not test any other AI tool, and our percentages should not be read as a measure of any other product. What carries over is the kinds of error, which match what researchers have documented in large language models generally.
- The review was AI-assisted, not a teacher panel. See the method and its limits.
- Every error we found has been fixed, and the generator has been changed so it makes them less often.
What we found
We made 2,244 corrections across 450 of the 601 resources (74.9%). 151 resources needed nothing. The chart shows the share of all 601 with at least one problem of each type; one resource can appear in several rows.
| Type of error | Resources | % of 601 | Distinct issues | Corrections |
|---|---|---|---|---|
| Wrong answer keys and scoring guides | 98 | 16.3% | 134 | 149 |
| Math and science errors | 115 | 19.1% | 244 | 258 |
| Factual errors in other subjects (history, geography, English, RE...) | 122 | 20.3% | 260 | 266 |
| Invented statistics, studies, dates or details | 79 | 13.1% | 147 | 171 |
| Misquotes, or details that don't match the text | 24 | 4% | 39 | 40 |
| Wrong standards codes or exam details | 16 | 2.7% | 24 | 26 |
| Pitched at the wrong level for the age group | 88 | 14.6% | 294 | 578 |
| Wrong country: terms, units, money, spelling | 125 | 20.8% | 235 | 326 |
| Truncated, garbled or broken text | 82 | 13.6% | 100 | 146 |
| Consistency, timing, instructions and wording | 147 | 24.5% | 217 | 284 |
"Distinct issues" counts each written reason once per resource; one issue can need several edits ("corrections"), for example a wrong fact repeated on a slide, a worksheet and its answer key. A further 55 edits only kept other parts consistent with a fix and are not counted as errors.
What simple automated checks found across all 601
| Automated check | Resources flagged (before) | Flags (before) | Flags (after review) |
|---|---|---|---|
| Standards reference that isn't a real code | 83 of 601 | 185 | 0 |
| UK spelling in a US resource | 140 of 196 | 799 | 0 |
| UK school terms in a US resource (Year 7, GCSE, pounds...) | 48 of 196 | 111 | 1 |
| US spelling in a UK resource | 57 of 405 | 87 | 4 |
| Exam-style marks or timings for 4 to 7 year olds | 18 of 601 | 74 | 15 |
| Questions worth 3+ marks for 4 to 7 year olds | 32 of 601 | 32 | 0 |
143 of our 196 US resources (73%) had UK spellings or school terms before the review. The checks are blunt, so some flags are false alarms, for example a word quoted on purpose or a deliberately wrong statement in a true/false quiz.
Why AI gets lessons wrong
None of this is mysterious once you know how these models work. Below are five reasons, in plain English. Each is a known behaviour of large language models in general; the percentages and examples show how it turned up in our own library.
1. It predicts what sounds right. It doesn't look facts up.
A language model writes by predicting the next most likely words, based on patterns in the text it learned from. Unless a tool adds a separate fact-checking step, nothing in that process looks anything up. Most of the time likely words are true words; sometimes they aren't, and the wrong version reads just as smoothly. Researchers call this "hallucination", and a major survey notes that it is a particular concern precisely because the output is fluent.[3] Researchers at one AI developer compare it to a student guessing in an exam: "Like students facing hard exam questions, large language models sometimes guess when uncertain, producing plausible yet incorrect statements instead of admitting uncertainty."[2]
Why it matters: A core fact reversed, inside the "correction" students were meant to learn from.
Why it matters: Monroe was not President until 1817, and presidents do not sign amendments.
2. It fills gaps with plausible inventions.
When a lesson "needs" a statistic, a date, a study or a quotation and the model doesn't have a real one to hand, it tends to produce something with the right shape. Invented references are a well-documented example: a peer-reviewed study in Scientific Reports found that many of the academic citations generated by two versions of a popular chatbot did not exist at all.[4] In lessons, the same habit produces percentages, page numbers and quotations that look authoritative and have nothing behind them.
Why it matters: The statistic had no source. It looked authoritative and was made up.
Why it matters: The quotation is not in the novel. Students could have learned it and used it in an essay.
3. It doesn't check its own working.
An answer key is written the same way as everything else: as likely-sounding text. The model does not go back, do the sum, balance the equation or test each multiple-choice option. So the key can confidently say "A" when the answer is C, or call an unbalanced equation balanced. This is the error teachers are least likely to spot, because the key is the part we trust.
Why it matters: The key marked a wrong option as correct.
Why it matters: The key said the atoms matched. They didn't.
4. It blurs levels and countries.
Models learn from a huge mix of material: university textbooks and children's books, British and American sources, exam papers from many systems. Unless they are held firmly to one age and one country, they drift towards whatever they have seen most of. That shows up as A-level vocabulary in a high school lesson, exam language for 5 year olds, UK health guidance in a US lesson, and standards codesthat look official but don't exist. In our case this was made worse by our own older templates, which asked for exam-style structures at every age. That part was our design choice, not the model's, and we have changed it.
Why it matters: Exam-room instructions for children who are just learning to read and write.
Why it matters: There is no such code. It looked official, so it would be easy to copy into a plan.
Why it matters: The wrong country's guidance, and the figure was wrong for that guidance too.
5. Long documents lose track of themselves.
A full lesson pack is long: a plan, slides, a worksheet, an answer key, teacher notes. Each part can read perfectly on its own while disagreeing with another part: timings that don't add up, a story that changes between the slides and the worksheet, a definition the quiz then contradicts. Sentences can also stop half way, or a writing frame can lose its blanks.
Why it matters: Each part read well on its own; together they contradicted each other.
What to check in any AI-made resource
This applies to resources from any AI tool, including ours. AI can save real time, but the UK's Department for Education is clear that "Any content produced requires critical judgement to check for appropriateness and accuracy."[1] Seven checks, based on what we found:
- Work the answer key yourself. Do every question you plan to set, or at least every math and science one, before you trust the key. In our library, 16.3% of resources had a wrong key.
- Find every quotation in the text. AI can write lines that sound exactly like the author and are not in the book. If a quote matters, find it on the page.
- Treat precise numbers, studies and dates as unverified. "A 2018 study found..." or "90% of..." with no source should be checked or deleted.
- Look up standards codes. Check them against the official Common Core, NGSS or state documents. Codes that look official can be invented.
- Read it as your students would. Is it pitched for this grade? Watch for test-style points and command words ("evaluate") for young children, and for content from later grades.
- Check it's from your country. Spelling, money, units, organizations and guidance (USDA or NHS, dollars or pounds) should match your students' lives.
- Check the parts agree with each other. Timings that add up, a story that stays the same on the slides and the worksheet, sentences that finish, blanks that are really there.
What we changed in our generator
The library was fixed resource by resource. The bigger job was stopping the generator from making the same mistakes for every teacher. These changes went live on 3 October 2026, and each one targets one of the reasons above:
- Level ceilings (reason 4). Each age band has a written list of what is in and out of scope. Output is checked against it, and anything above the level is rewritten for that age.
- Verified standards codes (reason 4). Codes are checked against a reference list (2,030 US standards from Common Core, NGSS and the C3 Framework, plus 76 checked UK references and 60 exam-board qualifications). Anything not on the list is removed rather than shown.
- A quotation bank (reason 2). 98 verified quotations from 13 commonly taught texts and poems are offered to the generator, and quotes from public-domain texts can be checked against the full text.
- Answer checking (reason 3). Written sums are recalculated, chemical equations in answer keys are balanced, multiple-choice questions must have exactly one right option, and true/false answers must agree with their labels.
- Fact checks on knowledge organizers (reasons 1 and 2). Key dates, people and statistics are kept only if they can be verified; the rest are dropped or flagged.
- Region rules (reason 4). US resources get US spelling, terms, units and guidance; UK resources keep UK ones.
- Consistency fixes (reason 5). Cover-lesson timings are reconciled with the lesson's phases, and writing frames keep their blanks.
- Report a mistake. Resources made in LessonHQ now have a button to tell us about an error, and reports come straight to us.
Did it work? Our own before-and-after tests
We gave the previous and the new generator the same prompts and compared the results. Each row has its own scale: the amber bar is the previous generator, the indigo bar is the new one.
We also read 30 previous/new pairs side by side in full. The new generator was judged better in 18, about the same in 11 and worse in 1 (an early-years counting deck that stopped at six, fixed afterwards).
Read these numbers with care. The samples are small (30 and 26 prompts). The checks measure our own rules, so a drop to zero means "none of the problems this check looks for", not "no errors". The side-by-side judgments were made by AI reviewers working to a written rubric, not by teachers. AI output changes every time it runs, and random slips still happen with the new generator, for example the odd flawed question on a math worksheetfor young children. The fixes reduce the errors; they don't remove the need to check.
Method
- What was checked. All 601 resources in LessonHQ's official library on 2 October 2026: 405 for UK schools and 196 for US schools, across 11 types (lesson packs, slide decks, worksheets, lesson plans, knowledge organizers, exit tickets, activities, cover lessons, homework booklets, revision sheets and feedback sheets). All were made by earlier versions of the LessonHQ generator. A copy of every resource was saved before any change.
- Step 1: automated checks over all 601: standards codes against verified lists, region spelling and terms, written sums, multiple-choice keys, true/false labels, cut-off text and exam-style marks for young children.
- Step 2: a line-by-line read of every resource (132 on 2 October, the other 469 on 3 October). The reading was AI-assisted: AI reviewers, a different model from the ones that wrote the resources, read each resource in full against a written checklist (facts, answer keys, calculations, quotations, sources, level, region, consistency). Where a fact was in doubt it was checked against published sources. Every correction was recorded with a written reason and applied by script, so each one can be traced and reversed.
- Categories. Each correction was put in one category by fixed rules applied to its written reason (first match wins). Factual errors in math, science and computing resources count as "Math and science errors". A random sample of 100 categorised corrections was re-checked by hand and 90 were in the category a person would choose, so treat the category split as approximate. The totals do not depend on it.
- Counting. "Resources affected" means at least one issue of that type, out of all 601. The "error of fact, answer key, calculation, invented detail or quotation" figure combines the first five categories in the table. All numbers on this page are computed by a script from the saved review logs and the before/after copies, not typed in.
- Why AI gets it wrong. The five reasons are general explanations drawn from published research (sources below) and matched to our categories by us. They explain the kinds of error; they are not a measurement of any tool other than ours.
Limitations
- It describes one library made by one generator, before the changes above. It does not measure other AI tools, and it does not measure our current generator (the tests above do that, on a small scale).
- An AI-assisted review can miss errors and can make them. It was not peer-reviewed, and no panel of teachers checked every correction. The true error rate is more likely to be higher than lower.
- Some categories involve judgment, especially "wrong level" and "wording". Errors range from serious (a reversed science fact) to small (a missing word).
Sources
- Department for Education (2025). Generative artificial intelligence (AI) in education. GOV.UK policy paper. www.gov.uk/government/publications/generative-artificial-intelligence-in-education/generative-artificial-intelligence-ai-in-education
- Kalai, A. T., Nachum, O., Vempala, S. S. and Zhang, E. (2025). Why Language Models Hallucinate. arXiv:2509.04664. arxiv.org/abs/2509.04664
- Ji, Z. et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12), article 248. doi.org/10.1145/3571730
- Walters, W. H. and Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 13, 14045. www.nature.com/articles/s41598-023-41032-5
Data, citing and press
Download the summary data (CSV), free to reuse with credit (CC BY 4.0). Please cite as: Poyser, A. (2026). AI lesson tools aren't as accurate as you think. LessonHQ. For interviews or the full method, see our press page. How LessonHQ uses AI is explained on AI safety.
Aaron Poyser is the founder of LessonHQ. He isn't a teacher, but he's surrounded by them, and he builds tools that give teachers their time back.