When readers encounter a stranger in a story, they appear to file that person away by occupation rather than by who they are as an individual. This finding emerges from three preregistered online studies conducted by Yale researchers, each involving 300 American adults who read short narratives about everyday transactions while their reading speed was measured. The experiments tracked how quickly participants processed text, revealing that job titles perform substantial cognitive work in shaping expectations about what strangers will do and what they might know.
Participants read passages one phrase at a time on the platform Prolific, pressing the space bar to reveal each new segment. The researchers were not concerned with the content of the complaints or requests in the stories; rather, they focused on who was receiving them. In one scenario, a blocked car belonged to a postal worker, and later in the story, a different postal worker appeared. In another version, the same situation involved customers instead. The wording of the grievance remained identical across versions, only the job titles changed.
The results showed a striking difference. When readers encountered the final words of passages involving customer-to-customer substitution, they took approximately 213 milliseconds longer than when reading about postal-worker-to-postal-worker substitution. The model estimate underlying the significance test suggested an even larger gap of about 262 milliseconds. This research was published in Communications Psychology on 19 August by Aaron Baker, Yarrow Dunham, and Julian Jara-Ettinger of Yale University. Their central claim is that job titles enable people to anticipate stranger behavior, form expectations about what that stranger knows, and treat one job-holder as readily replaceable by another—a process fast enough to show up in reading times.
When the replacement is another postal worker
The third study began interactions with one person and concluded with someone else in two of its three conditions: a pen borrowed from one agent and returned to another at a doctor's office, a computer borrowed from one hotel employee with its password requested from a second, and the blocked car scenario at the post office.
When a postal worker was replaced by another postal worker, the final region took 1352.24 milliseconds to read. When the same customer remained throughout the interaction, it took 1327.88 milliseconds. When one customer was replaced by a different customer, the time climbed to 1565.66 milliseconds. This last figure exceeded the average elsewhere in the study by more than 50 percent, where region-sized chunks typically took 1024.69 milliseconds. The authors drew this comparison specifically for the customer swap but not for the other two conditions.
The researchers concluded that readers were "equally unsurprised" by a same-role substitution and by no substitution at all, but surprised when the replacement person had no job to assume. No significance test compared the postal-worker swap directly against the unchanged-customer condition, and these two differ in both the role itself and whether any swap occurred. The authors offered an explanation: a salesperson can step away mid-transaction and another can take over, because the role is what the interaction depends on.
An exploratory analysis flagged in the preregistration found that 47 percent of the 293 participants who remained after excluding five who paused for over ten seconds and two with missing data had their slowest reading in the swapped-customer condition, compared to 29 percent for the swapped postal worker and 25 percent for no swap. With three conditions, random chance would sit at 33 percent, placing one cell above it and two below.
Reading as a stopwatch
The method is deliberately narrow, and that narrowness serves a purpose. In self-paced reading, a passage breaks into regions revealed one at a time, and the interval between key presses registers as processing difficulty. During reading itself, participants face no task beyond reading, so any slowdown emerges as a by-product rather than a conscious judgment. Only afterward were they asked recall questions and asked to rate how surprising or confusing each passage had been—the item combined both words, which matters for an argument distinguishing between difficulty and judgment.
The ratings aligned with reading times across all three studies. In the first, the most common response was "not surprising at all" for a role-consistent story, "pretty surprising" when the role-holder violated the script, and "very surprising"—the highest point on the scale—when the correct object was taken by someone with no role at all. In the third study, the most common response in the swapped-customer condition was still "not at all," at 38.3 percent, but nearly matched by "a little" at 35 percent, and this condition differed significantly from both comparisons.
Much of the experimental design focused on control conditions. A slowdown when a mechanic takes a cell phone off the counter could simply reflect that an auto shop primes car keys rather than phones, so a third condition kept the object constant and changed the person: a customer taking car keys took 1068.43 milliseconds, the mechanic taking the phone took 1020.98 milliseconds, and the mechanic taking the keys took 849.34 milliseconds. The paper tested the first two against the last but not against each other. A slowdown when a stranger mentions your ulcer could simply reflect surprise that a non-staff character spoke at all, so another condition had that character mention their own ulcer instead.
Most procedural matters were routine. Six reading times went missing due to a software error in each of the second and third studies. The reading-time data violated normality assumptions, and the authors report that results survived a log transform. The model used for surprise ratings in the first study violated a proportional-odds assumption; they refitted it with the assumption relaxed, found the result held, and still report the preregistered model.
One issue stands out as non-routine and deserves attention alongside the second study's result. The authors had preregistered a rule to "exclude all reading times outside +/- 2 standard deviations of that participant's mean reading time." They later reconsidered this as a mistake, and their reasoning is sound: the rule removes each reader's slowest responses, while the prediction being tested is precisely that certain conditions make readers slow down. Their supplementary figure shows it removing the most data from exactly the conditions where a slowdown was predicted, "precisely because our experiment worked as predicted." The main text therefore excludes only pauses over ten seconds—12 of the 2,700 readings across all three studies. Under the preregistered rule, however, the second study's key contrast vanishes: the fellow patient's remark comes out "no different" from the doctor's, at p = 0.83. The first study's one control contrast also falls short of significance, at p = 0.17, though its main contrast holds at p = 0.017. Outlier handling in the third study was never preregistered at all. None of this remains hidden—the paper flags the deviation in its methods and sets the entire matter out in a supplementary note. It does mean the second part of the argument rests on an analysis choice made after the authors had seen their own data.
A doctor mentioning your ulcer
The second study shifted focus from actions to knowledge. A doctor in a waiting room remarking on the ulcer in your stomach took 896.47 milliseconds, which the authors describe as comparable to the 809.09 milliseconds that region-sized chunks took elsewhere in that study. A fellow patient making the identical remark pushed it to 1044.55 milliseconds. A patient remarking on the ulcer in their own stomach did not, on average, appear to slow readers down.
This third condition allows the authors to argue against a simpler explanation, though they remain cautious about the rest. It remains possible, they write, that readers "were not explicitly representing knowledge, and only forming a low-level expectation about the kinds of utterances people produce"—a doctor expected to say doctor-ish things, with no conscious thought about what the doctor knows. Even so, they add, roles could still support that thought later, when someone takes time to form it.
The errors ran one way
These recall analyses were exploratory: the authors label them as such in the first study and describe them as explorations in the other two, despite a methods section stating all analyses were preregistered. The abstract nonetheless presents the misreporting result as a headline. Readers answered two questions about each story, each a choice between two options, so random chance sits at 50 percent. No condition approached it. What varied was the direction of the errors. In the first study, 92 percent correctly named the role-holder who took the object, against 69 percent when the person doing it had no job in the scene—a customer in two stories and a friend in the third. In the second, 97 percent named the doctor as the speaker, against 74 percent when the same line came from another patient.
The paper's account of this is a hypothesis rather than a finding, and the key point belongs in the quotation. Because readers could not scroll back, the authors suggest, they "might have therefore assumed they misread the passages and reconstructed the events in a role-consistent way. If so, the errors were not failures of memory, but of reconstruction."
A separate literature points the other direction, and the paper cites it: people are found elsewhere to remember unexpected events better, not worse. A reconciliation is proposed but not tested, and this design cannot resolve it.
What the labels leave out
This account was written from the paper itself rather than from a summary, which is the only expertise available here; the sharpest cautions come from the authors themselves. Every story named its roles explicitly, in words. A physical setting offers uniforms and context instead, and the authors note that whether those cues produce the same rapid expectations is a question for future work. The participants were all in the United States, working with jobs that read as jobs in the United States, and the authors list cultural variation first among their limitations. They also decline the most convenient available headline: "our work does not imply that role-based reasoning is faster than Theory of Mind."
The article carries the publisher's early-access banner. It is an accepted manuscript, and the version of record will replace it.
Two further questions remain open, and the authors identify both. One concerns how fast these expectations actually form and which cue triggers them; the paper raises the possibility that they are pre-loaded as we enter a familiar space, before any person appears. The other concerns how far the library of roles extends. A cashier can be replaced mid-transaction and a friend cannot, though the authors allow that a friend might be replaced over a longer timeframe, and they ask for a more precise map of what lies between. The materials, data, and analysis code are posted, which makes both questions cheaper to pursue than to debate.
Source: Silicon Canals



