A pair of economists engaged by the European Research Council's executive agency examined scoring patterns across a vast dataset of grant applications. Their analysis of 37,073 proposals submitted between 2019 and 2023 reveals that when reviewers and applicants shared both citizenship and country of employment, the first-stage evaluation score increased by 0.164 points on the ERC's five-point scale. The connection to actual funding outcomes remains less direct and depends on correlation rather than causation.

According to findings published on 13 August in Nature Human Behaviour, Europe's premier research funding body shows measurable bias toward applicants holding the same passport as their reviewers and working in the same nation. The phenomenon reflects not differences among applicants themselves, but rather how different reviewers evaluate identical proposals.

Nicola Fuchs-Schündeln and Felix Holub, researchers at the WZB Berlin Social Science Center, examined all applications to the European Research Council during the specified period. When an applicant and panel member shared both citizenship and country of work, the reviewer's first-stage score climbed by 0.164 points—equivalent to 16 percent of a standard deviation. By comparison, shared gender produced a score increase of just 0.024 points.

Methodology and research design

The ERC distributes proposals among 27 panels, each containing approximately 15 senior researchers whose identities remain confidential until after the evaluation period concludes, with the exception of the panel chair. During the initial stage, panelists independently rate a condensed version of each assigned proposal on a scale from 1 to 5 in half-point increments, providing separate evaluations for both the project and the principal investigator. The panel subsequently convenes to assign each proposal a grade of A, B, or C; only A-rated proposals advance.

The analytical approach allowed the researchers to control for two variables simultaneously. A proposal fixed effect captures all characteristics of the proposal itself, including unmeasured quality dimensions. A reviewer fixed effect accounts for individual variation in scoring stringency. The remaining comparison isolates the effect of interest: how the same proposal receives different scores from reviewers who match the applicant versus those who do not.

This methodological framework permits the authors to characterize the individual-score findings as causal rather than merely correlational—a significant distinction from earlier research that primarily compared success rates rather than examining scores assigned to identical material.

Citizenship and country-of-work effects

Shared citizenship alone, without matching country of work, raised first-stage project scores by 0.075 points (95 percent confidence interval 0.048 to 0.101). Matching country of work without citizenship produced a 0.056-point increase. When both factors aligned, the boost reached 0.164 points. Scores for evaluating the principal investigator followed similar patterns with somewhat smaller magnitudes, rising 0.083 points when both citizenship and country matched. Second-stage results showed comparable effect sizes.

The authors note that single-attribute matches produced effects "on average two to three times larger than for gender." When comparing the combined-match coefficient directly to the gender effect, the ratio approaches seven.

The researchers investigated potential innocent explanations for these patterns. If matched reviewers simply possessed aligned research interests, controlling for intellectual proximity should reduce the effect. They adjusted for semantic similarity between each proposal and the reviewer's 20 most-cited papers, for direct and indirect coauthorship during the previous decade, and for whether the proposal's ERC keyword matched the most frequently used keyword among that reviewer's assessments that year. The coefficients remained essentially unchanged.

Another hypothesis suggested that matched reviewers possessed superior knowledge of applicants, which should produce more extreme scores in both directions. Instead, matched reviewers increased the likelihood of top scores without increasing the likelihood of bottom scores. The probability of a perfect 5.0 on a first-stage project score rose by 0.014 when citizenship alone matched, against a baseline of 0.047, while the probability of a 1.0 score actually declined slightly. This pattern contradicts what better information typically produces.

The authors examined whether shared institutional prestige created advantages, classifying the ten institutions receiving the most grants in each panel as elite. They found no statistically significant evidence that reviewers favored applicants from their own institutional tier.

Panel-level decisions show weaker patterns

Individual scores represent inputs to a larger process. Funding decisions emerge from the panel's collective judgment, and at this stage the evidence transforms in character. Each proposal receives only one panel verdict per stage, the fixed-effect design becomes unavailable, and findings become correlational rather than causal. The authors acknowledge this distinction explicitly.

Correlational associations do appear at the first stage. One additional panel member sharing both citizenship and country correlated with a 1.7 percentage point higher likelihood of an A grade, compared to a baseline of 32.6 percent. A chair sharing both attributes correlated with a 3.1 percentage point increase, though this result reached only marginal significance at a p value of 0.055. At the second stage, neither A grades nor funding decisions showed significant associations with reviewer-applicant matching.

Across both stages combined, the authors estimate that one additional doubly-matched panel member correlates with a 1.0 percentage point higher probability of funding, relative to a baseline funding rate of 14.2 percent.

The most concrete disparity involves five countries that collectively submit just over half of all proposals: Germany, the UK, France, Italy, and Spain. Applicants from these nations encounter panels where 4.6 percent of members match both citizenship and country of work, compared to 1.5 percent for applicants from all other countries. These applicants also face higher proportions of members matching country only (4.5 versus 1.8 percent) and citizenship only (3.8 versus 2.6 percent). When the authors apply these three differences through their model, the implied funding-probability advantage reaches 1.7 points versus 0.8, a gap representing approximately 6 percent of the baseline success rate.

Gender effects prove minimal

The gender finding—likely to be oversimplified in subsequent discussion—does show that same-gender reviewer-applicant pairs produced higher individual scores, and these estimates achieved statistical significance. However, the effects ranged from 2.4 to 4.7 percent of the relevant standard deviations, and the authors explicitly state that effects of this magnitude are "unlikely to influence final decisions." Additionally, effects varied by academic field, appearing in life and physical sciences but failing to reach significance in social sciences and humanities.

At the panel level, the gender result not only disappears but reverses direction. The coefficient on the share of same-gender panel members in second-stage funding decisions was negative, with a p value of 0.081. This represents a marginal result on a correlational estimate and should not be interpreted as evidence that same-gender panels disadvantage applicants. It does, however, contradict what readers might expect from headlines emphasizing gender bias in grant review.

Funding source and data access constraints

The authors' disclosures carry significance equal to their coefficients. Fuchs-Schündeln and Holub state explicitly that they worked as contracted experts for the ERC Executive Agency, that their research represents "a direct product of this expert engagement," and that they received fees plus travel allowances and reimbursement. Fuchs-Schündeln, who holds a professorship at Goethe University Frankfurt, has also been on the applicant side of the process: she received an ERC Starting Grant in 2010 and a €1.6 million Consolidator Grant announced in November 2018.

The access enabling this study simultaneously prevents replication. Application-level data remain confidential under the authors' expert contract and cannot be published. Any researcher wishing to reproduce the analysis must petition a specific ERCEA unit, with access granted at the agency's discretion. The authors have released their code with synthetic data, allowing the analytical pipeline to run, which represents progress but does not constitute true replication.

Three additional limitations appear within the paper itself. Panel-level results represent associations only, so an unobserved pattern—such as certain panels attracting stronger proposals from particular countries while also seating more members from those nations—could generate the observed correlations. The individual-score effects constitute averages of fractions of a single scoring increment, and the paper does not demonstrate a single grant changing hands due to these biases. Furthermore, no statistical correction was applied for multiple hypothesis testing, a limitation the authors disclose.

As of 21 August 2026, eight days following publication, no correction or formal comment had appeared on the article page. The journal lists Matthias Egger and Alexander Petersen among its peer reviewers and has published their reports.

Proposed procedural reforms

The authors' recommendations focus on procedures rather than moral judgments. They suggest assigning reviewers to limit same-country and same-citizenship pairings, following the model used by World Trade Organization dispute panels that exclude panelists from involved countries. Diversifying panels would mechanically reduce matching frequency.

Alternatively, reforms could target the scoring process itself. A second assessment could evaluate proposals near the decision threshold, where the authors argue small score differences carry greatest consequence. Proposals could be scored on materials stripping the investigator's identity, with the investigator evaluated separately. The authors also mention borrowing partial lottery systems that others have proposed for marginal cases.

The European Research Council, established in 2007 and operating a budget exceeding €16 billion from 2021 to 2027, evaluates proposals according to declared scientific excellence. The paper frames homophily as working against this goal while systematically advantaging applicants from large, well-represented research systems and disadvantaging those from smaller nations. The first stage represents the critical juncture: proposals failing to advance at this stage never reach interviews, external referees, or panel discussions where an initial misreading might be corrected.

Source: Silicon Canals