arXiv:2412.17803cs.LG2024-12中稿 · IEEE ICHI 2025被引 9

研究医疗语言模型在标签不平衡下的表现与公平性问题

Examining Imbalance Effects on Performance and Demographic Fairness of Clinical Language Models

  • 分析性别、年龄、种族等维度的数据分布不均对模型影响
  • 发现多数类特征相似性比数据量更影响模型性能
  • 适合关注医疗AI公平性的研究人员参考

数据不平衡是生物医学应用中语言模型面临的核心挑战,尤其在ICD编码预测任务中,标签和人口统计分布极不均衡。尽管先进语言模型被广泛应用于生物医学领域,但很少有研究系统考察数据不平衡对模型性能及不同人群公平性的影响。本研究通过主流生物医学语言模型,在标准基准数据上分析了性别、年龄、种族及健康社会决定因素层面的不平衡问题。结合多种评估指标与统计分析,探究数据不平衡对性能差异与群体公平性的效应。结果表明,数据不平衡显著影响模型表现与公平性,但多数类特征相似性可能是更关键因素。该研究为开发更具公平性与鲁棒性的医疗语言模型提供了重要洞见。

原文摘要 · Abstract (English)

Data imbalance is a fundamental challenge in applying language models to biomedical applications, particularly in ICD code prediction tasks where label and demographic distributions are uneven. While state-of-the-art language models have been increasingly adopted in biomedical tasks, few studies have systematically examined how data imbalance affects model performance and fairness across demographic groups. This study fills the gap by statistically probing the relationship between data imbalance and model performance in ICD code prediction. We analyze imbalances in a standard benchmark data across gender, age, ethnicity, and social determinants of health by state-of-the-art biomedical language models. By deploying diverse performance metrics and statistical analyses, we explore the influence of data imbalance on performance variations and demographic fairness. Our study shows that data imbalance significantly impacts model performance and fairness, but feature similarity to the majority class may be a more critical factor. We believe this study provides valuable insights for developing more equitable and robust language models in healthcare applications.

医疗AI公平性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。