Researchers staged the largest human–AI creativity comparison to date in January 2026, pitting several language models against responses from 100,000 participants on a simple but revealing exercise. The findings complicate the narrative of machine dominance: GPT-4 achieved a higher average score than the typical human participant, yet the distribution of human performance tells a different story, with the upper half of respondents consistently outscoring all tested models.
The Divergent Association Task: A Measurable Creativity Benchmark
The test, known as the Divergent Association Task (DAT), asks participants to list ten English nouns that are as semantically distant from one another as possible. Rather than opposites or items from a single category, the goal is to find words whose meanings and uses diverge maximally. The challenge becomes apparent quickly: once "dog" is written, "cat" feels inevitable; "ocean" naturally suggests "mountain."
Published in Scientific Reports on January 21, 2026, the peer-reviewed study employed a scoring system that converts each valid word into a numerical representation based on its contextual usage across large language corpora. The algorithm then calculates semantic distances among the first seven valid words, producing 21 possible pairs. The mean distance of these pairs, multiplied by 100, yields the final DAT score.
The methodology builds on work introduced by Jay Olson and colleagues in a 2021 paper in the Proceedings of the National Academy of Sciences. That earlier research, conducted with 8,914 participants, demonstrated moderate to strong correlations between DAT performance and established creativity measures, including tasks requiring unusual applications for everyday objects or conceptual bridges between disparate ideas.
Importantly, the researchers acknowledge that the DAT measures divergent verbal association—one recognized component of creative cognition—rather than creativity in its entirety. It does not capture the full arc from initial idea to finished invention, compelling narrative, functional design, or enduring artwork. The task provides a standardized, automated scoring mechanism that applies equally to human and machine responses, offering more rigor than subjective evaluation of whether a poem "feels inspired."
The Sample: Large, Balanced, but Not Universal
The researchers selected 100,000 English-speaking participants from a larger dataset, deliberately balanced across gender (50% men, 50% women) and age. The cohort was divided into five equal groups spanning ages 18–29, 30–39, 40–49, 50–59, and 60 or older. Most participants came from the United States, with smaller contingents from the United Kingdom, Canada, Australia, and New Zealand.
All participants had accessed the public DAT website after learning about it through news coverage, social media, or word of mouth. This volunteer recruitment method produced an unusually large and deliberately balanced dataset, but not a random global sample. The study's reference to "humanity" represents this English-speaking volunteer population rather than a comprehensive census of human creative capacity.
For the machine side, researchers tested OpenAI's GPT-3.5, GPT-4, and GPT-4-turbo; Anthropic's Claude 3; Google's GeminiPro; and several open-source models including Pythia, StableLM, RedPajama, and Vicuna. An expanded comparison included systems released between January 2023 and June 2025. Each model generated 500 responses, with researchers initiating a fresh conversation for each iteration to prevent prior answers from influencing subsequent ones.
GPT-4 Beats the Average, But Not the Top Performers
In the primary comparison, GPT-4 achieved the highest average score among all tested models and exceeded the overall human mean by a statistically significant margin. GeminiPro's average proved statistically indistinguishable from the human mean. Notably, GPT-4-turbo performed worse than the older GPT-4 on this task—a cautionary finding against assuming that newer or more computationally efficient models automatically demonstrate greater divergence.
However, this headline-friendly result obscures a more complex picture. Five hundred GPT-4 generations produced a higher mean score than the mean of 100,000 human responses, but the distribution of human scores reveals a striking pattern when examined beyond the average.
When researchers sorted human DAT scores and constructed benchmarks from the upper 50%, upper 25%, and top 10%, the pattern shifted decisively. The mean score of the most creative half of the human sample exceeded the mean of every language model in the comparison. The upper quartile scored even higher, and the top decile demonstrated the clearest separation.
This distinction matters substantially. The result does not mean that each of the 50,000 people in the upper half outperformed every single one of the model's 500 responses; some model outputs overlapped with or exceeded individual high human scores. Rather, the comparison examined group centers, not individual contests. Additionally, the study's model set concludes in June 2025, so the findings describe where those particular systems performed under those specific prompts and settings, not the state of all available systems in August 2026.
Machine Consistency Versus Human Diversity
Beneath the averages lies a revealing pattern about how language models approach the task. Across separate GPT-4 sessions, the word "microscope" appeared in approximately 70% of response sets, while "elephant" appeared in roughly 60%. GPT-4-turbo proved even more repetitive, with "ocean" surfacing in more than 90% of its sets.
Human participants exhibited strikingly different behavior. Their most frequent words were "car" at 1.4%, "dog" at 1.2%, and "tree" at 1.0%—far lower concentrations than any model.
This repetition does not contradict GPT-4's high average score. The DAT rewards semantic distance within individual responses but does not directly penalize similarity across multiple responses. A model can discover that microscope, elephant, and several other words form a reliably distant set, then reuse those ingredients across different answer sets. Each individual list may span considerable semantic space while the population of lists follows a predictable route.
Humans, collectively, generated a far broader array of routes. Their average individual list was less divergent than GPT-4's, but their choices across the sample were less concentrated. This distinction carries practical implications: a system can excel at producing an unexpectedly varied answer on demand while still tending toward similar answers when many users pose the same request. Individual novelty and collective diversity are not synonymous.
Temperature and Prompting Shape Model Performance
The team investigated whether GPT-4's score could be modified without retraining the model. Temperature—a parameter controlling how heavily generation favors the most probable next token—proved influential. Lower temperature settings produce more deterministic responses, while higher settings permit less likely continuations, typically increasing variation though sometimes at the cost of coherence.
GPT-4's mean DAT score rose significantly with temperature. At the highest tested setting of 1.5, its mean reached 85.6, surpassing 72% of human scores. Repeated words also became less frequent as temperature increased.
Prompting strategy similarly affected results. When researchers tested instructions encouraging different linguistic approaches, asking GPT-4 to consider etymology—the origins and structures of words—produced the highest mean score among tested strategies.
These findings underscore that claims about "a model's creativity" remain conditional on multiple variables: Which model? Which version? Which prompt? Which temperature? How many samples? Who selects the final output? These are not mere technical footnotes but rather descriptions of the creative system under evaluation. A human who adjusts the prompt, requests additional candidates, and recognizes which output merits attention becomes part of the process generating the result.
Beyond Word Lists: Testing Other Formats
The researchers extended their analysis beyond noun lists, applying related automated measures to haiku, movie synopses, and flash fiction. GPT-4 scored above GPT-3.5 across these writing formats on their measure of semantic divergence, while human-written haiku and synopsis samples retained advantages on key comparisons.
These exercises provided supporting evidence that the DAT captures something relevant beyond isolated word lists. However, they did not transform semantic distance into a comprehensive theory of artistic creation. Brilliant work can emerge from semantically close words, while technically distant combinations can prove useless, incoherent, or merely peculiar. Genuine creativity typically demands both novelty and some form of fit: usefulness in engineering, insight in science, emotional resonance in fiction, or coherence within aesthetic choices.
The authors acknowledge these limitations. Automated measures cannot fully capture usefulness, convergent thinking, or expert judgment. An ideal future benchmark would combine computational scoring with human evaluation and test models on tasks kept private until assessment.
One concern warrants particular attention: the DAT prompt has been public since 2021. Because commercial model training data remain opaque, researchers could not exclude the possibility of prior exposure. Models may have learned task examples or discussions of scoring strategies during training. While the team's public code and data repository enhance transparency, they cannot reveal what exists within proprietary training corpora.
The 100,000 human participants also came with limited metadata. Researchers lacked information about occupations or creative experience. It remains plausible that writers, musicians, editors, and other practiced creators concentrated in the upper tail, but the study could not test this hypothesis.
Implications for Creative Work and Human Judgment
The study illuminates why language models feel creative in everyday use. They generate remote associations swiftly, consistently, and at negligible marginal cost. For someone confronting a blank page, this capacity offers genuine utility.
Yet ideation represents only one stage of creative work. The full process also demands deciding what is appropriate, recognizing when an unusual connection contains genuine insight, rejecting fluent but empty output, developing promising fragments, and revising until the result coheres within a larger whole.
The upper human tail matters because creative industries do not always prioritize average ideas. Their value frequently resides in rare responses and in the judgment to identify them.
The model repetition finding introduces another reason to maintain human involvement. If many users rely on the same system with similar prompts, each may receive something feeling fresh in isolation while broader culture gradually becomes more homogeneous.
The most defensible conclusion proves narrower, and more useful, than declaring an outright winner. On this specific test of divergent verbal association, GPT-4 exceeded the human average. The most creative half of the sample maintained a higher average than every tested model, and the top 10% extended the gap further.
Machines have grown remarkably effective at raising the floor of ideation. The ceiling still depends on uncommon human divergence, and on the distinctly human capacity to recognize which unexpected association deserves development into something substantial.
Source: Silicon Canals



