Scientific authorship · AJPH · 2013–2022
Who submits.
Who gets accepted.
A ten-year audit asks what name-based algorithms can reveal about disparities in scientific publishing, and what they can miss.
Important
These are predicted categories, not self-identified identities.
The algorithms infer the likely ethnic origin of names and binary gender. The study uses those probabilities as proxies, with explicit uncertainty.
01 · The study
One journal. Ten years. A full submission pipeline.
Corresponding authors only · United States submissions · Final editorial decisions · Multiple imputation to retain prediction uncertainty
02 · Race and ethnicity
Submission volume and acceptance tell different stories.
Share of submissions
n = 17,667Acceptance rate
Predicted categoriesPredicted White authors had the highest acceptance rate. Predicted Asian authors had the lowest.
03 · Gender
More submissions from women. A lower acceptance rate.
Predicted women
57.4%of submissionsPredicted men
42.6%of submissionsThe pattern of lower acceptance for women appeared across most predicted racial and ethnic groups. Asian women were the exception, with a rate slightly above Asian men.
04 · Intersections
The widest contrast appears when categories are combined.
White men
Black men
White women
Hispanic men
Black women
Hispanic women
Asian women
Asian men
Acceptance rates for corresponding US authors, using pyethnicity and Gender API classifications.
THE TAKEAWAY
Measure inequity. Measure the measurement. Then collect better data.
Algorithms can help journals examine historical patterns when demographic data are missing. They should complement, not replace, voluntary self-identified data and careful investigation of structural conditions.
Read the article in the American Journal of Public Health ↗