让领域专家参与数据生成,有效降低医疗AI的代表偏差。
Explanatory Debiasing: Involving Domain Experts in the Data Generation Process to Mitigate Representation Bias in AI Systems
- 通过设计通用指南,引导领域专家参与数据构建。
- 35名医疗专家参与实证研究,证明可降偏差且不损失准确率。
- 适合需高可靠性、强解释性的医疗等专业领域AI开发。
代表性偏差是人工智能系统中最常见的偏差类型之一,导致模型在少数群体数据上表现不佳。尽管从业者采用多种方法缓解此类偏差,但其效果常受限于去偏过程中的领域知识不足。本文提出一套通用设计指南,以有效整合领域专家参与代表性去偏。我们在医疗场景中实现该指南,并通过包含35名医疗专家的混合方法用户研究进行评估。结果表明,引入领域专家可在不损害模型准确率的前提下显著降低代表性偏差。基于研究发现,我们为开发者提供可落地的建议,推动基于通用设计指南构建更稳健的去偏系统,确保领域专家在去偏过程中被有效纳入。
原文摘要 · Abstract (English)
Representation bias is one of the most common types of biases in artificial intelligence (AI) systems, causing AI models to perform poorly on underrepresented data segments. Although AI practitioners use various methods to reduce representation bias, their effectiveness is often constrained by insufficient domain knowledge in the debiasing process. To address this gap, this paper introduces a set of generic design guidelines for effectively involving domain experts in representation debiasing. We instantiated our proposed guidelines in a healthcare-focused application and evaluated them through a comprehensive mixed-methods user study with 35 healthcare experts. Our findings show that involving domain experts can reduce representation bias without compromising model accuracy. Based on our findings, we also offer recommendations for developers to build robust debiasing systems guided by our generic design guidelines, ensuring more effective inclusion of domain experts in the debiasing process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。