提出预测代表性框架,揭示皮肤癌AI模型对深肤色人群的系统性偏差。
Predictive Representativity: Uncovering Racial Bias in AI-based Skin Cancer Detection
- 以预测结果为核心,评估模型在不同肤色群体中的公平性表现。
- 在HAM10000和哥伦比亚临床数据集上均发现深肤色者检测性能显著更低。
- 适合关注医疗AI公平性、数据正义与算法审计的研究者阅读。
人工智能系统日益参与医疗决策,但算法偏见与不平等结果仍存,尤其对历史边缘化群体。本文提出预测代表性(Predictive Representativity, PR)框架,将公平性审计焦点从数据集构成转向结果层面的公平性。通过皮肤病学案例研究,我们评估了基于广泛使用的HAM10000数据集及哥伦比亚独立临床数据集(BOSQUE Test set)训练的AI皮肤癌分类器。分析显示,按肤色分组存在显著性能差异,即使数据集按比例采样,深肤色个体的分类器表现始终较差。我们认为,代表性不应视为数据集的静态属性,而应是模型预测在子群体与部署场景中动态、情境敏感的特性。PR通过量化模型在不同子群体与场景中公平性的泛化能力来实现这一转变,并提出外部可迁移性准则以明确公平性泛化的阈值。研究强调事后公平性审计、数据集文档透明化与包容性模型验证流程的伦理必要性。本工作提供了一种可扩展工具,用于诊断AI系统中的结构性不公,推动对数据驱动医疗中公平性、可解释性与数据正义的讨论。
原文摘要 · Abstract (English)
Artificial intelligence (AI) systems increasingly inform medical decision-making, yet concerns about algorithmic bias and inequitable outcomes persist, particularly for historically marginalized populations. This paper introduces the concept of Predictive Representativity (PR), a framework of fairness auditing that shifts the focus from the composition of the data set to outcomes-level equity. Through a case study in dermatology, we evaluated AI-based skin cancer classifiers trained on the widely used HAM10000 dataset and on an independent clinical dataset (BOSQUE Test set) from Colombia. Our analysis reveals substantial performance disparities by skin phototype, with classifiers consistently underperforming for individuals with darker skin, despite proportional sampling in the source data. We argue that representativity must be understood not as a static feature of datasets but as a dynamic, context-sensitive property of model predictions. PR operationalizes this shift by quantifying how reliably models generalize fairness across subpopulations and deployment contexts. We further propose an External Transportability Criterion that formalizes the thresholds for fairness generalization. Our findings highlight the ethical imperative for post-hoc fairness auditing, transparency in dataset documentation, and inclusive model validation pipelines. This work offers a scalable tool for diagnosing structural inequities in AI systems, contributing to discussions on equity, interpretability, and data justice and fostering a critical re-evaluation of fairness in data-driven healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。