The first major global study of student AI use has landed, and the headline finding is uncomfortable for anyone building AI tools for learning. Students who use AI chatbots daily for schoolwork score worse than those who do not, and the gap is not small.
The OECD's 2025 Programme for International Student Assessment, released on September 8, tested more than 760,000 15-year-olds across 91 countries. It is the first PISA cycle conducted since generative AI went mainstream. Among students who never or almost never use AI to draft writing assignments, the average science score was 509. Among students who use AI for that purpose every day or almost every day, it was 481. That 28-point gap, adjusted for socio-economic status, equals roughly a year and a half of teaching.
Only 14% of students reported never using AI for schoolwork. The technology is already embedded in how the next generation learns. The question is whether it helps them think or replaces the thinking entirely.
What did the PISA 2025 results actually find?
The OECD's data is the most comprehensive look yet at AI use in education, and it resists a simple verdict on whether the technology is good or bad for learning.
The clearest signal is negative. Students who use AI to perform specific tasks that substitute for their own cognitive work, like summarizing assigned reading or drafting texts for writing assignments, perform relatively worse. The gap is sharpest for daily users: a 28-point deficit in science compared to non-users, after adjusting for socio-economic status.
Frequency matters in a non-linear way. Students who used AI only once or twice a year and those who used it daily both tended to score worst. Monthly and weekly users landed in between, with weekly users occasionally outperforming everyone, including non-users. The pattern suggests a bell curve: too little AI use to matter at one end, over-reliance at the other, and a useful middle ground.
Country-level adoption varies enormously. Over 95% of Vietnamese students reported using AI tools, compared to roughly 60% in Japan. Almost-daily use remains relatively uncommon globally at less than 20%, but the direction of travel is clear. Nearly half of students in developed economies say they use AI chatbots to help them learn at least weekly, according to AP's coverage of the findings. The chart below shows the flip side: the share of students who never use AI, ranging from about 40% in Japan to just 4% in Vietnam.

The broader PISA results are grim reading even before you get to the AI question. Reading scores across OECD countries fell by 28 points between 2015 and 2025, the steepest decline on record since data collection began in 2000. Mathematics dropped 22 points over the same period. The share of OECD 15-year-olds who are low performers across all three subjects, reading, math, and science, rose from 16% in 2022 to 20% in 2025. Reuters reported that teenagers are reading at the worst levels seen this century, with increased screen time and falling rates of reading for enjoyment coinciding with the decline.
Is the problem AI itself or how students use it?
The OECD is careful not to claim that AI causes lower scores. The data is correlational, and the report acknowledges that the relationship between AI use and performance is complex. A student who is already struggling may reach for AI more often, not the other way around. But the pattern is consistent enough, and the sample large enough, that the findings demand attention.
The report draws a critical distinction between two modes of AI use: using AI to do the thinking, and using AI to support the thinking. OECD Director for Education and Skills Andreas Schleicher put it bluntly: "In the same way that we do not become fit by watching sports but by doing sports, learning does not occur through the consumption of content, but as a productive cognitive struggle of the mind with new material." He called for AI to function as a "scaffold, not a crutch."
The data backs this up. Students who use AI to conduct preliminary research on a new topic, or who use it broadly to "help me learn," saw less of a performance drop than those using it to draft or summarize. And there is a genuinely positive signal: among students who said they use AI weekly to help them learn, scores slightly exceeded those of non-users.
The most actionable finding is about what happens when students are taught to evaluate AI output. Among daily AI users, those who are regularly asked to assess the quality of AI-generated information in their lessons scored 13 points higher in science than daily users who were not, equivalent to more than half a year of teaching. The training did not eliminate the gap with non-users entirely, but it narrowed it significantly.
What does this mean if you are building AI tools for education?
If you build edtech, learning features, or any AI product where users are supposed to acquire knowledge or skill, the PISA data is a warning about design defaults. The evidence points to a clear divide between tools that replace cognitive work and tools that scaffold it.
- Products that replace effort will hurt your users' outcomes. The 28-point gap for daily users who draft and summarize with AI is the market telling you that cognitive offloading has a cost. If your tool's core value proposition is "we do the work for you," you are building the thing the OECD says correlates with worse learning.
- Products that scaffold effort have a defensible position. Weekly users who engage with AI as a learning aid outperform non-users. Features that prompt the user to think, check, or evaluate, rather than just consume, align with the pattern that works.
- Critical assessment training is a feature, not an afterthought. The 13-point boost from teaching students to evaluate AI output is one of the most concrete numbers in the report. If your product includes AI-generated content, building in prompts, checks, or exercises that make the user assess that content could be the difference between a tool that helps and one that harms.
The business consequence is straightforward. Edtech markets are regulated and reputation-sensitive. Independent research on how people use AI in practice shows that real usage patterns often diverge from intended design. The PISA data gives regulators and school districts their first large-scale evidence base for restricting or conditioning AI use. Expect more districts and ministries to ask what your product does to promote productive use rather than substitution.
The curiosity data is worth noting too. Curiosity levels correlate with AI use, peaking among daily users. That means the students most drawn to AI tools are also the most engaged. The challenge for product designers is to channel that curiosity into learning rather than having it consumed by the tool.
What should builders and education leaders do now?
The PISA findings suggest a few concrete moves for anyone building or deploying AI in learning environments.
- Build for the weekly user, not the daily user. The data shows weekly, moderate use is where AI helps. Design for sessions that supplement learning rather than replace it. Limit features that encourage continuous drafting or summarizing.
- Make AI output assessment a core interaction. The 13-point science boost from critical assessment training is the single most actionable finding in the report. If your product surfaces AI-generated content, pair it with questions, checks, or reflection prompts that force the user to evaluate what the AI produced.
- Distinguish between learning modes and performance modes. A tool that helps a student learn a concept is different from one that helps them produce an assignment. The PISA data suggests the former can help and the latter can hurt. Products that conflate the two are the most risky.
- Watch for regulatory movement. The OECD is the closest thing education policy has to a global standard-setter. When it says AI use correlates with worse outcomes and calls for targeted, scaffolded use, that language will show up in national and local policy. The UK's science scores actually rose in this cycle. Bloomberg's coverage notes that Schleicher pointed to China, Japan, Singapore, and Estonia as countries that have succeeded in using AI to support learning without overusing it.
There are caveats. The data is correlational and self-reported. Students self-select into AI use patterns for reasons the study controls for imperfectly. The PISA cycle runs every three years, so the next chance to measure longitudinal effects is 2028. And the 2025 results capture a moment when AI tools were still relatively new in education; usage patterns and tool capabilities will both evolve.
The design question nobody can afford to get wrong
The PISA 2025 data is the first hard evidence that the defaults of AI tools matter enormously for cognitive outcomes. When the tool does the thinking, the user gets worse at thinking. When the tool supports the thinking, the user can get better. That is a design decision, and it is now measurable at a scale of 760,000 students across 91 countries.
For builders, the finding is that badly designed AI in education is bad, and the evidence for that is strong enough to shape policy. The products that will win in this market are the ones that make the user do the work, with AI as the scaffold. The OECD just handed you the spec.
Sources
- OECD - PISA 2025 Results (Volume I)
- Bloomberg - AI Use in Classrooms Linked to Lower Test Scores, OECD Data Suggests
- AP News - Student scores in high-income countries hit a low point, test shows
- Reuters - Teen reading slumps to worst this century due to surge in screen time
- The Star - School students who use AI get worse test scores, OECD warns
- The Verge - Students who use AI generally score worse at school
