arXiv:2604.14514cs.AIcs.CE2026-04中稿 · publication in the…

biomedical AI模型训练数据存在严重种族偏见,可能加剧医疗不公。

Perspective on Bias in Biomedical AI: Preventing Downstream Healthcare Disparities

  • 分析4514篇基因组论文,仅2.7%报告族裔信息
  • 主流数据集如CellxGene和GEO中欧洲血统样本占绝对主导
  • 提出可溯源、开放、评估透明三大原则应对早期偏见

医疗不平等现象持续存在,常归因于筛查、诊断与治疗的获取差异。然而,本文指出关键偏见可能更早出现——在数据收集与研究优先级设定阶段,尤其是在分子与组学数据研究中。尽管大量研究聚焦于组学数据采集,但相关人口统计信息常缺失:对2015至2024年4514篇PubMed收录的组学论文进行自动化分析发现,仅有2.7%报告族裔或血统信息,地理来源报告率仅为2.5%。对常用模型训练数据集(如CellxGene和GEO)的分析显示,欧洲血统数据占据主导地位。随着生物医学基础模型成为发现核心,其通过大规模预训练并复用于多种下游任务,可能放大早期偏见,导致难以逆转的系统性不平等。为此,我们倡导社区遵循三项基本原则:可溯源性(Provenance)、开放性(Openness)与评估透明性(Reliability through Evaluation Transparency),以提升偏见可见度,支持更明智的模型开发与部署决策。

原文摘要 · Abstract (English)

Healthcare disparities persist across socioeconomic boundaries, often attributed to unequal access to screening, diagnostics, and therapeutics. However, this perspective highlights that critical biases can emerge much earlier, during data collection and research prioritization, long before clinical implementation, particularly in studies focused on molecular and omics data. A vast number of studies focus on collecting omics data, but the demographic information associated with these datasets is often not reported, and when it is reported, it reveals substantial biases. An automated analysis of 4514 PubMed-indexed omics publications from 2015 to 2024, examining reporting across multiple demographic dimensions, reveals limited reporting overall; for example, only 2.7% of studies report ancestry or ethnicity information and geographic origin reporting is limited to 2.5%. Analysis of large-scale datasets commonly used for model training, such as CellxGene and GEO, reveals substantial population bias where European-ancestry data dominates. As biomedical foundation models become central to biomedical discovery with a paradigm in which base models are pretrained on large datasets and reusing them repeatedly for many different downstream tasks, they risk perpetuating or amplifying these early-stage biases, leading to cascading inequities that regulatory interventions cannot fully reverse. We propose a community-wide focus on three foundational principles: Provenance, Openness, and Reliability through Evaluation Transparency. Together, these principles can help make biases and limitations more visible to model developers and users, supporting more informed model development, evaluation, and deployment decisions in biomedical AI.

AI偏见医疗公平组学数据基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。