Researchers in Greece conducted a preregistered trial in which teachers evaluated student work already carrying a mark from someone else. The study manipulated two variables: whether teachers were told the existing mark came from a human colleague or from an algorithmic system, and whether that mark was too lenient or too harsh. The attribution shifted teacher behaviour only in one direction—when the mark was unfairly low—and only among certain groups.

The experimental setup was straightforward. Teachers accessed a survey containing a student exercise with five components, each marked right or wrong. A mark of five out of ten was already assigned. The correct answer, worth eight points based on the exercise's own logic, sat plainly on the screen. Teachers then entered their own mark.

Testing the human-in-the-loop promise

Policy frameworks for algorithmic decision-making typically assume that human oversight catches machine errors. A system proposes, a professional reviews, the professional corrects mistakes. This logic underpins approaches to algorithmic grading, triage and screening. Sofoklis Goulas, Rigissa Megalokonomou and Panagiotis Sotirakopoulos tested this assumption with 1,339 active teachers across Greece, publishing their findings in PNAS Nexus in June 2026.

The experimental design

Two dimensions varied across the study. The source of the mark changed—either a colleague or an algorithmic system. The direction of error also changed. All exercises matched the teacher's subject area, and within each subject all teachers saw identical work. When four of five answers were correct, the fair mark was eight and the suggested five was unfairly harsh; when one of five answers were correct, the fair mark was two and the same five was unfairly generous. The suggested mark itself never varied from five.

Researchers measured what they termed the grading fairness gap: how far each teacher's mark deviated from the objectively correct one. Each participant evaluated a single exercise and assigned one mark, contributing one data point. The trial was preregistered with the AEA registry.

The scope was deliberately narrow. One country. One labelled exercise per teacher. Surveys took a median of 7.2 minutes to complete, not a full marking session. The study included only day schools, excluding evening, vocational and special-education contexts. Researchers recorded marks without asking teachers about their reasoning until afterwards.

When the mark was unfairly harsh and attributed to a human source, teachers' average gap from the correct grade was 1.384 points. Despite having the right answer just one subtraction away on the screen, teachers anchored heavily to the mark they received, regardless of its stated origin. This anchoring effect itself was substantial before any label effect entered.

Only harsh marks showed the label effect

When the harsh mark was attributed to an algorithm, the gap widened to 1.584 points. The difference of 0.300 points proved statistically significant (p = 0.003 with controls, p = 0.022 without), representing a 22 per cent increase relative to the human baseline.

The generous mark told a different story. Under the human label, the gap was 1.930 points; under the algorithm label, 1.813 points. This difference of −0.108 was not statistically distinguishable from zero (p = 0.685). The label produced no detectable effect when the initial mark was unfairly lenient.

Minor discrepancies appear between the paper's prose and its tables. On the harsh version, text reports 1.386 human and 1.586 algorithm while the table shows 1.384 and 1.584; on the generous version, text gives 1.818 algorithm and 1.942 human against table values of 1.813 and 1.930. The direction of differences remains consistent across both presentations, and the 22 per cent figure holds either way.

Under the human label, the generous version produced a larger gap (1.930) than the harsh version (1.384), but the paper does not directly test this comparison. Teachers deviated asymmetrically: they softened unfairly harsh marks and inflated unfairly generous ones. The same upward adjustment moved marks toward fairness on the harsh side while pushing them away from fairness on the generous side. The label effect operated on top of this existing asymmetry rather than replacing it.

Severity as a signal of competence

After marking, teachers rated the source on five dimensions: ability, comprehension, fairness, intent and responsibility. Across both versions, teachers rated the algorithmic source substantially lower than the human one, with the widest gaps appearing on the generous version. Notably, the perceived ability deficit was smaller when the mark was harsh. The researchers interpreted strictness itself as evidence that the grader possessed expertise.

Mediation analysis explored this mechanism. On the harsh version, perceived ability accounted for an indirect effect of 0.218 (p < 0.001) and responsibility for 0.143 (p = 0.016), while comprehension, fairness and intent showed no significant paths. On the generous version, all five perception channels ran in the opposite direction and reached significance, yet the total effect remained indistinguishable from zero.

The perception questions were asked after teachers learned the mark's source, making these findings suggestive rather than definitively causal. They also came after the grading decision itself—a sequencing choice designed to prevent attitude questions from influencing the mark. This order leaves room for an alternative interpretation: a teacher who has just accepted a mark may rate its source favourably to justify that decision. The shares attributed to ability and responsibility—roughly 73 per cent and 47 per cent respectively—represent each path's coefficient divided by the total 0.300 effect. Their sum exceeds 100 per cent because overlapping paths are reported separately, a limitation the authors acknowledge.

The effect concentrated among tech-confident teachers

Deference to the algorithmic label was not uniform. On the harsh version, it reached statistical significance for teachers under 51 years old (0.463), those holding a master's or doctorate (0.442), humanities specialists (0.533), and those rating their own technological literacy highly (0.486). For older teachers, those with only a bachelor's degree, STEM specialists and those with low tech literacy, the effect was small and not statistically significant. On the generous version, no subgroup showed a significant difference.

The authors note an important caveat: confidence intervals for some group comparisons overlap, so differences between groups themselves remain unestablished. The effect concentrated among people most comfortable with technology—the opposite of what a narrative about wary teachers would predict.

The sample composition skewed in particular directions. Eight per cent held a doctorate compared to roughly 2 per cent of Greek K-12 teachers overall, aligning with where the effect appeared; average age was 49 against 40 for the profession, moving the opposite way. The authors conclude their findings may apply most directly to relatively experienced teachers in the Greek system.

Scepticism about AI grading ran deep

End-of-survey questions revealed limited enthusiasm for algorithmic grading. On a scale from −5 to +5, teachers' average belief that AI could grade fairly was 0.03; their willingness to allow it, −1.03; their sense that doing so would be ethical, −1.33. Nearly half used generative tools at least weekly for lesson planning, yet more than half rarely or never encouraged colleagues to adopt them.

An open-ended question at the survey's close invited additional comments. Among those expressing reservations, the authors' coding found three-quarters focused on moral or ethical concerns—the ill student, the child whose home circumstances should factor into the mark—and one-quarter on technical limitations. The paper does not specify how many comments were submitted overall. The scepticism was genuine and articulated. On the harsh mark, where all five perception channels worked against the algorithm and the label left no net effect, this scepticism appears to have done its work; on the harsh mark, where the total effect ran the opposite direction, it did not.

What the study can and cannot show

The task's simplicity was intentional and shapes what the findings mean. Making the correct mark deducible from a list of right and wrong answers isolates the label effect. It also means the study measures reluctance to override a source under ideal conditions—no ambiguity, no fatigue, no stack of forty scripts with a deadline. The authors acknowledge this and note that real systems arrive with explanations, confidence scores and accuracy records that a vignette cannot replicate.

This study documents one controlled scenario: a vignette, a label attribution, and a resulting mark. It cannot speak to what happens in an actual marking session. Where the analysis ventures beyond the paper's explicit claims—regarding the mediation shares and why teachers might rate a grader favourably after accepting its mark—those interpretations are noted as such.

The teachers in this experiment had everything needed on screen and still landed well away from the fair mark. The algorithmic label detectably increased that distance only when the machine was being unjustly harsh. If this is what oversight resembles when the error sits one subtraction away and the mark belongs to no one, what value does it hold in a setting where errors are buried in complexity and the mark determines where a fifteen-year-old student proceeds next?

Source: Silicon Canals