Some behavioural measures of test engagement, which allow for trend comparisons and can therefore help contextualise the observed reading trends, can be derived from other domains assessed in PISA, and even from the questionnaire.
Given the earlier findings that indicated somewhat stronger declines in the latter part of the reading test, compared to the initial section, one may wonder whether this reflects more generally a drop in the capacity to sustain effort; and whether a two-hour computer-based test may not exceed the endurance of 15‑year‑olds. To what extent are the negative trends in reading reflective of shorter concentration spans?
Differences in performance by hour of testing and domain order
A first angle to investigate student fatigue consists in comparing the performance of students who sat the reading test in the first hour and in the second hour of testing. A similar comparison can be conducted for all subjects; and if the differences in performance are indicative of the students’ capacity to sustain a cognitive effort, in general, results should be relatively consistent regardless of the subject. Indeed, by virtue of the random allocation of test booklets to students, the groups of students that are being compared in each case are strictly equivalent.
Tables I.A1.3, I.A1.4 and I.A1.5 present these comparisons for science, mathematics and reading. Performance differences between the first and second hour are often negative, and sometimes relatively large (more than 10 score points in reading, on average across 35 OECD countries, regardless of the year). However, the between-country variation in these differences is not strongly correlated across subjects, and the differences themselves do not seem to have widened over time: more often than not, the negative trend in reading between 2018 and 2025 was even stronger for students who took reading as the first subject (Table I.A1.4). Only a few countries/economies stand out for showing a more negative reading trend for students who took the reading test in the second hour (Belgium, Kosovo and Lithuania) than for students who took the reading test in the first hour.
An alternative interpretation of these differences emphasises the role of interactions across subjects. In this interpretation, performance is not only influenced by a domain’s position in the test, but by the precise nature of what students did in the hour before. For example, doing a test that is perceived as difficult, by most students, in the first hour (e.g. mathematics) may have a negative effect on students’ reading performance in the second hour; while doing an easier test in the first hour may have a positive effect. On average across OECD countries domain-order effects are not pronounced (Table I.A1.6); but this average does not necessarily reflect the absence of domain-order effect: it may result from the large variety of forms that domain-order effects take across countries and economies. For example, in Chinese Taipei, students who took the reading test after a mathematics test achieved higher scores than students who took it after the assessment of “learning in a digital world”; while a difference in the opposite direction is observed for students in Finland.
Differences in performance between the beginning and the end of the science test
A second angle to investigate the role of fatigue, and how it may have impacted performance trends (and reading trends in particular), exploits the science test, and in particular its linear version, which was used for 25% of test-takers. The item bank in science was divided in three item sets (A, B and C), from which sequences of about 11 to 13 items were extracted (these sequences are referred to as “testlets”). Each linear test consisted of three such testlets, and the order in which these testlets were administered was rotated across students, such that, out of these 25% of student, one third started with item set A, one third with item set B, and one third with item set C; and similarly, for the second and third testlets. The change in the percentage of correct responses over the three possible positions of an item set in the linear test therefore provides an indirect measure of test fatigue.
Table I.A1.2 shows that in most countries, results declined modestly over the course of the science test; only Korea and Kenya exhibit the opposite pattern of better results on the same items when they are located at the end of the test, compared to the beginning. The largest declines, with differences in P+ exceeding 4 percentage points, were observed in Bulgaria, Israel and Thailand (in descending order of the magnitude of the decline).
Overall, student fatigue during the PISA test, in general, does not appear to play a significant role in the observed trends; countries where fatigue effects are more pronounced in 2025 did show similar declines, compared to 3 or 7 years earlier, as countries in which fatigue effects are less pronounced.
Patterned response behaviour (straightlining)
A third indicator of fatigue, or perhaps more generally of disengagement in the PISA test, is based on mixed scales in the student questionnaires. Mixed scales are item batteries that measure a single construct using both positively worded and negatively worded items. An example – and the only mixed scale that can be compared over three cycles of PISA, from PISA 2018 to PISA 2025 – is the sense-of-belonging scale, which contains apparently contradictory statements such as “I feel lonely at school” and “I make friends easily at school”. In both cases, students are asked to indicate the extent to which they agree with the sentence provided.
Students who do not pay attention to the sentence, or to the response scale, may sometimes respond with the same response option (e.g. “strongly agree”) regardless of the item; thus contradicting themselves half of the time (the sense-of-belonging scale contains exactly three positively worded items and three negatively worded items). Students are said to “straightline” because their responses to these items are all perfectly aligned, on the response sheet – and most likely misaligned with how they truly feel. Straightlining is not the only possible pattern of disengagement, but is perhaps the easiest to detect.
On average across 23 OECD countries, the percentage of students who “straightlined” in the sense-of-belonging item battery was somewhat lower in 2025 than in both 2022 and 2018. This suggests that students were not less engaged with the PISA test and questionnaire than in previous years. In fact, there are only limited exceptions where straightlining increased significantly, and only when considering the most extreme version of it (when students always select the “strongly agree” or “strongly disagree” option), which remains very rare. The largest increase, from 0.9% to 2.2% (+1.3 percentage points), was observed in Latvia.