Research

AI lesson tools aren't as accurate as you think. We checked 601 AI-made teaching resources to find out why.

By Aaron Poyser, founder of LessonHQDownload the data (CSV)

AI writes like an expert. The sentences are fluent, the tone is confident, the layout looks finished. That is exactly what makes its mistakes easy to miss. Large language models are known to produce text that sounds right and isn't: the UK's Department for Education warns that generative AI "can produce nonsensical, inaccurate or false information presented as fact, known as hallucination".[1]

We wanted to know what that looks like in real lessons, so we started with our own. We read every one of the 601 resources in LessonHQ's official library, all written by earlier versions of our own AI generator, and logged every error. This page shows what we found, why it happens, what to check in any AI-made resource, and what we changed.

53%of our 601 resources had at least one error of fact, answer key, calculation, invented detail or quotation (320 resources)
16%had a wrong answer key or scoring guide (98 resources)
13%contained an invented statistic, study, date or detail (79 resources)
14%cited a standards code that doesn't exist (83 resources)

Whose mistakes these are

  • Every number on this page comes from our own library: 601 resources (405 UK, 196 US) made by earlier versions of our own generator. We are publishing our mistakes, not anyone else's.
  • We did not test any other AI tool, and our percentages should not be read as a measure of any other product. What carries over is the kinds of error, which match what researchers have documented in large language models generally.
  • The review was AI-assisted, not a teacher panel. See the method and its limits.
  • Every error we found has been fixed, and the generator has been changed so it makes them less often.

What we found

We made 2,244 corrections across 450 of the 601 resources (74.9%). 151 resources needed nothing. The chart shows the share of all 601 with at least one problem of each type; one resource can appear in several rows.

Share of our 601 AI-made resources with at least one error of each type
Wrong answer keys and scoring guides16.3%
Math and science errors19.1%
Factual errors in other subjects (history, geography, English, RE...)20.3%
Invented statistics, studies, dates or details13.1%
Misquotes, or details that don't match the text4%
Wrong standards codes or exam details2.7%
Pitched at the wrong level for the age group14.6%
Wrong country: terms, units, money, spelling20.8%
Truncated, garbled or broken text13.6%
Consistency, timing, instructions and wording24.5%
Hover or tap a row for counts. Source: LessonHQ library review, 2-4 October 2026. Full numbers in the table below and the CSV.
Type of errorResources% of 601Distinct issuesCorrections
Wrong answer keys and scoring guides9816.3%134149
Math and science errors11519.1%244258
Factual errors in other subjects (history, geography, English, RE...)12220.3%260266
Invented statistics, studies, dates or details7913.1%147171
Misquotes, or details that don't match the text244%3940
Wrong standards codes or exam details162.7%2426
Pitched at the wrong level for the age group8814.6%294578
Wrong country: terms, units, money, spelling12520.8%235326
Truncated, garbled or broken text8213.6%100146
Consistency, timing, instructions and wording14724.5%217284

"Distinct issues" counts each written reason once per resource; one issue can need several edits ("corrections"), for example a wrong fact repeated on a slide, a worksheet and its answer key. A further 55 edits only kept other parts consistent with a fix and are not counted as errors.

What simple automated checks found across all 601

Automated checkResources flagged (before)Flags (before)Flags (after review)
Standards reference that isn't a real code83 of 6011850
UK spelling in a US resource140 of 1967990
UK school terms in a US resource (Year 7, GCSE, pounds...)48 of 1961111
US spelling in a UK resource57 of 405874
Exam-style marks or timings for 4 to 7 year olds18 of 6017415
Questions worth 3+ marks for 4 to 7 year olds32 of 601320

143 of our 196 US resources (73%) had UK spellings or school terms before the review. The checks are blunt, so some flags are false alarms, for example a word quoted on purpose or a deliberately wrong statement in a true/false quiz.

Why AI gets lessons wrong

None of this is mysterious once you know how these models work. Below are five reasons, in plain English. Each is a known behaviour of large language models in general; the percentages and examples show how it turned up in our own library.

1. It predicts what sounds right. It doesn't look facts up.

A language model writes by predicting the next most likely words, based on patterns in the text it learned from. Unless a tool adds a separate fact-checking step, nothing in that process looks anything up. Most of the time likely words are true words; sometimes they aren't, and the wrong version reads just as smoothly. Researchers call this "hallucination", and a major survey notes that it is a particular concern precisely because the output is fluent.[3] Researchers at one AI developer compare it to a student guessing in an exam: "Like students facing hard exam questions, large language models sometimes guess when uncertain, producing plausible yet incorrect statements instead of admitting uncertainty."[2]

Math and science errors: 19.1%Factual errors in other subjects (history, geography, English, RE...): 20.3%
Real example from our library: UK, GCSE chemistry feedback sheet
Generated"In electrolysis the anode is the negative electrode where oxidation occurs."
CorrectedThe anode is the positive electrode, where oxidation occurs; the cathode is negative.

Why it matters: A core fact reversed, inside the "correction" students were meant to learn from.

Real example from our library: US, Grade 8 history knowledge organizer
Generated"James Monroe - signed the Bill of Rights into law as President in 1791."
CorrectedGeorge Mason - his Virginia Declaration of Rights influenced the Bill of Rights.

Why it matters: Monroe was not President until 1817, and presidents do not sign amendments.

2. It fills gaps with plausible inventions.

When a lesson "needs" a statistic, a date, a study or a quotation and the model doesn't have a real one to hand, it tends to produce something with the right shape. Invented references are a well-documented example: a peer-reviewed study in Scientific Reports found that many of the academic citations generated by two versions of a popular chatbot did not exist at all.[4] In lessons, the same habit produces percentages, page numbers and quotations that look authoritative and have nothing behind them.

Invented statistics, studies, dates or details: 13.1%Misquotes, or details that don't match the text: 4%
Real example from our library: UK, GCSE math slide
GeneratedBig statistic: "90% of architects use Pythagoras' Theorem".
CorrectedThe 3-4-5 right-angled triangle, a real and checkable example.

Why it matters: The statistic had no source. It looked authoritative and was made up.

Real example from our library: UK, A-level English literature homework
Generated"'He stretched out his arms toward the ______, but I think--' (Gatsby, Chapter 5)."
Corrected"he stretched out his arms toward the dark ______ in a curious way" (Nick describing Gatsby, Chapter 1).

Why it matters: The quotation is not in the novel. Students could have learned it and used it in an essay.

3. It doesn't check its own working.

An answer key is written the same way as everything else: as likely-sounding text. The model does not go back, do the sum, balance the equation or test each multiple-choice option. So the key can confidently say "A" when the answer is C, or call an unbalanced equation balanced. This is the error teachers are least likely to spot, because the key is the part we trust.

Wrong answer keys and scoring guides: 16.3%
Real example from our library: US, Grade 4 math quiz
Generated"Which fraction is larger?" A) 2/5 B) 3/8 C) 4/9 D) 5/12. Answer key: A.
Corrected4/9 is the largest. The options were rewritten so the keyed answer is right.

Why it matters: The key marked a wrong option as correct.

Real example from our library: US, Grade 10 chemistry answer key
Generated"Place coefficient 2 before Fe, 3 before O₂, giving 2Fe + 3O₂ → Fe₂O₃; atom counts match."
Corrected4Fe + 3O₂ → 2Fe₂O₃: 4 Fe and 6 O atoms on each side.

Why it matters: The key said the atoms matched. They didn't.

4. It blurs levels and countries.

Models learn from a huge mix of material: university textbooks and children's books, British and American sources, exam papers from many systems. Unless they are held firmly to one age and one country, they drift towards whatever they have seen most of. That shows up as A-level vocabulary in a high school lesson, exam language for 5 year olds, UK health guidance in a US lesson, and standards codesthat look official but don't exist. In our case this was made worse by our own older templates, which asked for exam-style structures at every age. That part was our design choice, not the model's, and we have changed it.

Pitched at the wrong level for the age group: 14.6%Wrong country: terms, units, money, spelling: 20.8%Invented standards codes (automated check): 13.8%
Real example from our library: UK, slides for 5 to 7 year olds
Generated"Answer in silence - no notes. 5 minutes."
Corrected"Think, then tell your partner. 5 minutes."

Why it matters: Exam-room instructions for children who are just learning to read and write.

Real example from our library: UK, elementary science lesson plan (ages 9-11)
GeneratedCurriculum link: "NCSS3-4b: The circulatory system".
CorrectedA verified statement from the National Curriculum program of study.

Why it matters: There is no such code. It looked official, so it would be easy to copy into a plan.

Real example from our library: US, Grade 9 health worksheet
Generated"The UK's Eatwell Guide recommends that around 30% of daily calories come from carbohydrates..."
CorrectedThe Dietary Guidelines for Americans recommend 45-65% of daily calories from carbohydrates.

Why it matters: The wrong country's guidance, and the figure was wrong for that guidance too.

5. Long documents lose track of themselves.

A full lesson pack is long: a plan, slides, a worksheet, an answer key, teacher notes. Each part can read perfectly on its own while disagreeing with another part: timings that don't add up, a story that changes between the slides and the worksheet, a definition the quiz then contradicts. Sentences can also stop half way, or a writing frame can lose its blanks.

Consistency, timing, instructions and wording: 24.5%Truncated, garbled or broken text: 13.6%
Real example from our library: UK, substitute lesson for ages 11 to 14 (math)
GeneratedOpening: "In this 60-minute lesson..." Task sheet: "You have 20 minutes."
CorrectedThe lesson's own timed phases added up to 50 minutes, and the main task phase was 30 minutes. Both lines were corrected to match.

Why it matters: Each part read well on its own; together they contradicted each other.

What to check in any AI-made resource

This applies to resources from any AI tool, including ours. AI can save real time, but the UK's Department for Education is clear that "Any content produced requires critical judgement to check for appropriateness and accuracy."[1] Seven checks, based on what we found:

  1. Work the answer key yourself. Do every question you plan to set, or at least every math and science one, before you trust the key. In our library, 16.3% of resources had a wrong key.
  2. Find every quotation in the text. AI can write lines that sound exactly like the author and are not in the book. If a quote matters, find it on the page.
  3. Treat precise numbers, studies and dates as unverified. "A 2018 study found..." or "90% of..." with no source should be checked or deleted.
  4. Look up standards codes. Check them against the official Common Core, NGSS or state documents. Codes that look official can be invented.
  5. Read it as your students would. Is it pitched for this grade? Watch for test-style points and command words ("evaluate") for young children, and for content from later grades.
  6. Check it's from your country. Spelling, money, units, organizations and guidance (USDA or NHS, dollars or pounds) should match your students' lives.
  7. Check the parts agree with each other. Timings that add up, a story that stays the same on the slides and the worksheet, sentences that finish, blanks that are really there.

What we changed in our generator

The library was fixed resource by resource. The bigger job was stopping the generator from making the same mistakes for every teacher. These changes went live on 3 October 2026, and each one targets one of the reasons above:

Did it work? Our own before-and-after tests

We gave the previous and the new generator the same prompts and compared the results. Each row has its own scale: the amber bar is the previous generator, the indigo bar is the new one.

Problems found by our checks: previous vs new generator, same prompts
Words and ideas above the level (e.g. A-level terms in a GCSE lesson)
78 paired generations (26 prompts x 3 rounds)
2442
Unverified dates, people or statistics in knowledge organizers
same 78 paired generations
430
Standards codes that don't exist
30 paired generations
160
UK spellings or terms in US resources
10 US paired generations
611
All automated-check flags
30 paired generations
10923

We also read 30 previous/new pairs side by side in full. The new generator was judged better in 18, about the same in 11 and worse in 1 (an early-years counting deck that stopped at six, fixed afterwards).

Read these numbers with care. The samples are small (30 and 26 prompts). The checks measure our own rules, so a drop to zero means "none of the problems this check looks for", not "no errors". The side-by-side judgments were made by AI reviewers working to a written rubric, not by teachers. AI output changes every time it runs, and random slips still happen with the new generator, for example the odd flawed question on a math worksheetfor young children. The fixes reduce the errors; they don't remove the need to check.

Method

Limitations

Sources

  1. Department for Education (2025). Generative artificial intelligence (AI) in education. GOV.UK policy paper. www.gov.uk/government/publications/generative-artificial-intelligence-in-education/generative-artificial-intelligence-ai-in-education
  2. Kalai, A. T., Nachum, O., Vempala, S. S. and Zhang, E. (2025). Why Language Models Hallucinate. arXiv:2509.04664. arxiv.org/abs/2509.04664
  3. Ji, Z. et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12), article 248. doi.org/10.1145/3571730
  4. Walters, W. H. and Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 13, 14045. www.nature.com/articles/s41598-023-41032-5

Data, citing and press

Download the summary data (CSV), free to reuse with credit (CC BY 4.0). Please cite as: Poyser, A. (2026). AI lesson tools aren't as accurate as you think. LessonHQ. For interviews or the full method, see our press page. How LessonHQ uses AI is explained on AI safety.

Aaron Poyser

Aaron Poyser is the founder of LessonHQ. He isn't a teacher, but he's surrounded by them, and he builds tools that give teachers their time back.