AI in education is software that adapts practice to each learner or generates explanations, feedback and speech on demand: AI tutors, writing feedback, lesson-planning assistants, captions and read-aloud. Randomized trials since 2023 show it can raise learning when it coaches students through problems, and lower it when students use it to get answers.

This guide covers what AI does in classrooms and training rooms today, what the peer-reviewed and independently evaluated trials found, the problems (cheating, unreliable AI detectors, over-reliance and bias), the guidance in force as of September 2026 in the US, England, Canada and the EU, student privacy under FERPA and COPPA, and an adoption plan for a school or a training team.

What AI does in classrooms today

Use is already common on both sides of the desk. In a Pew Research Center survey published in January 2025, 26% of US teens aged 13 to 17 said they had used ChatGPT for schoolwork, up from 13% in 2023. On the teaching side, a Gallup and Walton Family Foundation survey of 2,232 US public-school teachers, run in March and April 2025, found that six in ten had used an AI tool for their work during the 2024-25 school year, and 32% used one at least weekly.

The uses fall into five groups, and in each one some of the work should stay with a person:

UseWhat the AI doesWhat stays with the teacher
Tutoring and practiceExplains, asks guiding questions, gives hints, adjusts difficultyChoosing the content, checking the tutor is right
Feedback on writingComments on structure, argument and language against a rubricThe grade, and feedback that knows the student
Lesson planning and resourcesDrafts plans, quizzes, worksheets and versions for mixed levelsAccuracy and fit with the curriculum
AccessibilityLive captions, read-aloud, translation, syllable splittingDeciding who needs which support, and when
AdministrationDrafts emails, reports and lettersAny judgment about an individual student

Tutoring and study modes

AI tutors now come in two forms: purpose-built tutors inside learning platforms, and study modes inside general chatbots. Anthropic launched a learning mode with Claude for Education on April 2, 2025, OpenAI added study mode to ChatGPT on July 29, 2025, and Google added Guided Learning to Gemini on August 6, 2025. All three are designed to ask questions and give step-by-step guidance instead of the finished answer. OpenAI says study mode runs on custom system instructions and can behave inconsistently from one conversation to the next, so treat these modes as a helpful default, not a lock.

Feedback on writing

A model can comment on a draft in seconds, against the rubric the teacher supplies, so students can revise more often than a teacher could mark. Teachers rate it lower than their other uses: in the Gallup survey, 57% of teachers who used AI for grading and feedback said it improved the quality of that work, the lowest share of the nine tasks asked about (administrative work was highest, at 74%). England's Department for Education (DfE) says research shows generative AI could support feedback and tailored help, but that evidence on pupils using it themselves is still emerging.

Lesson planning

The firmest evidence so far is on teacher time. In a randomized trial evaluated by NFER for the Education Endowment Foundation (EEF), reported in December 2024, 259 science teachers in 68 secondary schools prepared Year 7 and Year 8 lessons either with ChatGPT and a guide, or without generative AI. The ChatGPT group spent 56.2 minutes a week on preparation against 81.5 minutes, a 31% saving, and an expert panel that did not know how each resource was made found no evidence of a difference in quality. A second EEF trial, of Oak National Academy's Aila lesson assistant in 86 primary schools, is due to report in autumn 2026. Gallup's teachers, self-reporting, estimated that weekly users save 5.9 hours a week.

Accessibility

Some of the most useful AI in a classroom is not generative at all. Speech recognition produces live captions for a lecture or video: Windows 11 live captions (version 22H2 and later) and Chrome's Live Caption both process audio on the device, so the audio never leaves the laptop (unless you turn on Chrome's Live Translate, which sends captions to Google). Microsoft's Immersive Reader, the reading tool in Word and OneNote that developers can also add to their own apps as an Azure service, reads text aloud, highlights parts of speech, splits words into syllables, shows pictures for common words and translates in real time. Microsoft designed it for new readers, language learners and people with dyslexia.

How personalized learning with AI works

Personalized learning is older than chatbots. An intelligent tutoring system (ITS) runs a loop: it diagnoses what the student knows from their answers, chooses the next problem or hint, gives feedback, and updates its model of the student. The evidence for that older generation is real but slower than the marketing suggests. In a RAND study published in 2014, schools in seven US states were matched in pairs and randomly assigned to Cognitive Tutor Algebra I, a personalized, mastery-based curriculum, or to their existing course. There was no effect in the first year; in the second year, results rose by about eight percentile points for the median student, a statistically significant gain in high schools but not in middle schools.

Large language models change two parts of that loop. They can explain in plain language, answer any question the student types and hold a conversation, which older tutors could not. They also lose two things older tutors had by design: a fixed, checked answer key and a reliable record of what the student has mastered. A language model predicts plausible text, and it can state a wrong solution fluently, a limit covered in where AI models fall short. The best-performing AI tutors in the research below put those parts back: their designers wrote the solutions, structured the steps and let the model do the talking.

A study assistant that answers only from material the instructor chose follows the same idea. Tools such as NotebookLM answer from the sources you add and cite them, which keeps a course's assistant tied to the course.

What the research says about AI tutors

The best evidence comes from randomized trials, where students are assigned to AI or no AI by chance. The main ones so far:

Study, publishedSettingDesignResult
Kestin et al., Scientific Reports, 2025Harvard introductory physics, fall 2023, 194 studentsCrossover: AI tutor at home vs the same lesson as active learningMedian learning gains more than double; median time 49 minutes against a 60-minute class
Bastani et al., PNAS, 2025A high school in Turkey, fall 2023, nearly 1,000 math studentsGPT-4 chat, GPT-4 tutor with safeguards, or no AI, during practicePractice grades up 48% (chat) and 127% (tutor); later exam without AI 17% lower for the chat group
De Simone et al., World Bank, 2025Nigeria, first-year senior secondary students, six weeksMicrosoft Copilot (GPT-4) for English, randomized0.31 standard deviations overall, 0.23 in English; largest gains for girls and stronger students
Wang et al., Tutor CoPilot, 2024900 tutors and 1,800 K-12 students in virtual math tutoring (US)AI suggests teaching moves to human tutors, randomized4 points more likely to master topics; 9 points for students of lower-rated tutors; about $20 per tutor a year
Pane et al., RAND, 2014Cognitive Tutor Algebra I, schools in seven statesPaired schools randomized, two school yearsNo effect in year one; about eight percentile points in year two

Read together, the trials say more about design than about AI. The Harvard tutor was written by the course's instructors, followed the same pedagogy as the class and was given worked solutions, because the team did not want to rely on GPT-4 to generate them. The Tutor CoPilot trial did not put AI in front of students at all: it coached novice tutors, who then asked more guiding questions and gave away fewer answers. The World Bank estimated its six-week program's gain as equal to 1.5 to 2 years of usual schooling, and the authors place it among the most cost-effective programs for improving learning outcomes.

The Turkish trial is the warning. Students with a plain chat interface to GPT-4 did much better on practice problems, then scored 17% lower than students who never had access once the tool was removed. The version with safeguards, whose prompt held teacher-written solutions and common mistakes and told it to give hints instead of answers, essentially removed that harm, though it did not raise exam scores either. The authors saw students using the plain version as a crutch during practice.

A tablet asks an AI box for help. A finished answer sheet stays behind a closed gate, and the AI sends back only a hint card, which leads up three steps to a worksheet the tablet completes itself.
Fig. 1 A tutor that gives the next step instead of the answer keeps the student doing the work that builds the skill.

The limits are worth stating. Most trials ran for weeks, not years, at one school or in one program, with tutors built by the researchers. None proves that any given commercial product will do the same in your school.

Over-reliance: when AI does the thinking

The mechanism behind the Turkish result has a name: cognitive offloading, handing recall, planning or reasoning to a tool. For an expert, that frees time. For a learner, the offloaded step is often the one the lesson was meant to build.

Governments and research funders are now designing for it. England's DfE updated its product safety standards for generative AI in education on January 19, 2026, adding standards on cognitive development. They expect learner-facing products not to give final answers, full solutions or complete worked examples by default, but to disclose help progressively, starting with hints, to ask the learner to attempt a step first, and to show a full solution only after a genuine attempt. Products should also report to teachers how often learners ask for this kind of offloading. In June 2026 the EEF opened a research fund of up to £2.5 million on how generative AI affects learning and cognition, with cognitive offloading as a central question.

In practice: give learners a tutor-style mode by default, keep the "just do it" assistant for staff and for tasks where the skill is not the point, and always measure what students can do without the tool.

Cheating, and why AI detectors are not proof

Generative AI broke one assumption: that an essay written at home shows the student can write it. The obvious fix, a detector, does not work well enough to decide a student's fate.

  • OpenAI withdrew its own detector. Its AI classifier, released in January 2023, caught 26% of AI-written text in OpenAI's own tests and wrongly flagged human text 9% of the time. It was withdrawn on July 20, 2023 because of its low accuracy.
  • Independent tests agree. Weber-Wulff and colleagues tested 14 detectors in 2023, including Turnitin and PlagiarismCheck. They found them neither accurate nor reliable, easily fooled by paraphrasing or machine translation, and concluded that the tools they tested should not be used in academic settings as evidence.
  • Detectors are biased against non-native writers. Liang and colleagues ran seven detectors on 91 TOEFL essays written by people and 88 US eighth-grade essays. The detectors were nearly perfect on the eighth-grade essays but flagged 61% of the TOEFL essays as AI-generated on average, and 89 of the 91 were flagged by at least one detector.

The reason is how detectors work. They score how predictable the word choices are, and plain, careful prose written in a second language is predictable. Text a student generated and then lightly reworded scores as less predictable, so it passes. The UK exam boards' Joint Council for Qualifications (JCQ) explains detectors the same way in its April 2025 guidance, notes that accuracy varies with factors including a student's English language competency, and treats a detector result as one piece of evidence to weigh with everything else the teacher knows about the student's work.

Two documents pass a detector with a dial. The one written with a pen is flagged with a warning mark, while the machine-written one, lightly edited with a pencil, passes with a check mark.
Fig. 2 A detector measures how predictable the words are, not who wrote them, so honest plain prose can fail while edited AI text passes.

Warning

Never sanction a student on a detector score alone. A false positive is an accusation against an honest student, and the studies above show it falls hardest on students writing in a second language.

What works better is designing assessment so the student's own thinking is visible:

  1. State the rules per task. Say which AI use is allowed (none, planning only, or full use with acknowledgement) and require students to disclose it, as JCQ requires for qualification work.
  2. Assess the process, not only the product. Collect outlines, drafts and version history, and give credit for the revision.
  3. Keep some writing supervised. Timed work in class, on managed devices where AI access can be switched off.
  4. Talk to the student. A five-minute conversation about their own argument reveals more than any score.

Bias, errors and equity

The same systems can treat students unequally in quieter ways. Detector bias is the documented case above. Generated content can also be wrong, out of date or biased, a risk England's DfE lists alongside hallucination, and a model can present a mistaken explanation with the same confidence as a correct one. Any AI output that reaches students needs the same check as a textbook page.

Gains can be uneven too. The World Bank trial found the largest effects for girls and for students with higher initial performance, which means a tool that helps everyone can still widen the gap between strong and weak students unless teachers target support. Tutor CoPilot showed the opposite pattern: the biggest benefit went to students whose tutors were rated lower. Access is a precondition as well: the DfE names connectivity among the barriers to effective use. Decisions about students, such as grades, placement or discipline, deserve the scrutiny set out in our guide to fairness in AI decision-making.

AI in education policy and guidance as of 2026

No single law governs classroom AI. What applies depends on where you teach:

WhereDocument and dateWhat it says
GlobalUNESCO, Guidance for generative AI in education and research (2023)A human-centred approach; governments should protect data privacy and set an age limit, with 13 as the minimum for independent use; institutions should validate tools
USExecutive Order 14277, April 23, 2025Promotes AI literacy; creates a White House Task Force on AI Education and a Presidential AI Challenge; prioritizes AI in teacher-training grants
USDepartment of Education Dear Colleague Letter, July 22, 2025Federal grant funds may pay for AI-based instructional materials, AI-enhanced high-impact tutoring and advising, with principles on privacy and parent involvement
EnglandDfE policy paper, updated August 12, 2025; product safety standards, updated January 19, 2026Fewer risks in teacher-facing use; pupil-facing use must meet data protection and safeguarding duties; no identifiable data in prompts
EnglandJCQ, AI Use in Assessments, April 30, 2025Work for qualifications must be the student's own; AI use must be acknowledged; misuse is malpractice
CanadaProvincial and territorial ministries; federal privacy commissioner's principles, December 7, 2023Education is run by the 10 provinces and 3 territories; the principles list education as a high-impact context and children as at particular risk
EUAI ActEmotion recognition in schools is banned since February 2, 2025; AI that decides access to education or scores exams is high-risk, with duties from December 2, 2027

Two things stand out. First, the US documents are about promotion and funding: they encourage AI use, AI literacy and teacher training rather than setting rules for classroom use. Second, the most specific rules for products now come from England's safety standards and the EU AI Act. The EU's high-risk duties for education were moved to December 2, 2027 by the AI Omnibus amendment, which entered into force on July 27, 2026.

Student privacy: FERPA, COPPA and their equivalents

Every prompt a student or teacher types can contain personal data: a name, a grade, an essay, a note about a disability. The question for any AI tool is who receives that data, for what purpose, and whether it is used to train the vendor's models.

  • FERPA (US). Education records are protected at schools and colleges that receive funds from US Department of Education programs. A school may share them without parental consent with a contractor only if that contractor counts as a "school official": it performs a service the school would otherwise use its own staff for, it is under the school's direct control for the use and maintenance of the records, and it follows FERPA's limits on redisclosure. In practice that means a signed agreement, not a teacher clicking through a consumer app's terms.
  • COPPA (US). Online services directed to children under 13 need verifiable parental consent before collecting their personal information. The FTC's guidance lets a school consent on parents' behalf only when the vendor collects data for the school's educational use and "for no other commercial purpose", and the vendor must still give the school notice and let it review and delete a child's data. The FTC finalized amendments to the COPPA Rule in January 2025; the amended Rule was published on April 22, 2025, with a year to comply with most changes. The FTC declined to finalize proposed changes for edtech companies operating in schools, so its FAQ on schools remains the guide.
  • England and the UK. UK GDPR and the Data Protection Act 2018 apply to schools and vendors alike. The ICO says an edtech provider falls under its Children's code when it uses children's data beyond the school's instructions, for example for product development, whatever its contract calls it. The DfE's standards expect suppliers to carry out a data protection impact assessment (DPIA) and not to use personal data for commercial purposes, including model training and fine-tuning, without a lawful basis.
  • Canada. Rules for schools are provincial and territorial, and the federal privacy commissioner's principles ask organizations using generative AI to give children particular protection.

Before signing, ask every AI vendor the same questions: what data it collects, whether any of it trains models, how long it keeps it and how deletion works, where it is processed, which subprocessors see it, how it handles users under 13, and how it tells you about a breach.

A school sends student records to a vendor's server through a padlocked contract. A dashed line from the vendor toward a model-training cluster is stopped by a barrier.
Fig. 3 A school can share student data with a vendor under contract; what the contract must stop is that data flowing on to train someone else's model.

Important

Staff should not paste identifiable student information into a personal AI account. England's DfE says data entered into generative AI tools should not identify anyone, and a personal account sits outside the agreement that FERPA's school-official route depends on.

How to adopt AI in a school or training team

The steps are the same for a school district, a university department or a corporate training team; only the laws in step 3 change.

  1. Start from a problem and a measure. More feedback on drafts in one year group's writing, or new support staff working on their own sooner, is a goal you can test. "Use AI" is not.
  2. Write a short AI-use policy. Which tools, for which tasks, by whom, what students or trainees must disclose, and who owns the decision to change it.
  3. Vet the tool before anyone signs up. Check the privacy terms against FERPA and COPPA, or run a DPIA in the UK, confirm that your data is excluded from model training, and prefer the institution-managed edition over personal accounts. UNESCO's guidance also expects institutions to validate a tool's educational fit, not only its security.
  4. Configure it for learning. Turn on tutor or study mode for learners, so hints come before answers, and choose tools that let teachers see usage, as England's standards expect.
  5. Train the staff who will use it. Cover what models get wrong, how to check output, and the fact that the professional who uses a draft remains responsible for it, as the DfE puts it.
  6. Redesign assessment before the pilot starts, using the four steps in the cheating section above.
  7. Pilot with a comparison group, and include one test taken without AI access. The Turkish trial shows why: practice scores alone would have made the harmful tool look like the best one.
  8. Review each term. Keep, change or stop the tool on the evidence, and record incidents such as wrong answers or data exposure.

For training teams, the most useful pattern is often an assistant that answers from your own course material and procedures, cites them and says when it does not know, alongside AI that coaches staff the way Tutor CoPilot coached tutors. Our AI and automation service builds assistants grounded in your own content that cite their sources and say when they do not know, with personal data removed before the model reads it. It starts by reviewing one real workflow with the people who run it, before anything is built.