研究视觉语言模型在医学影像诊断中因提示样本构成引发的种族差异
BiasICL: In-Context Learning and Demographic Biases of Vision Language Models
- 通过调整提示中的群体分布,发现模型会学习到不同人群的疾病基率
- 即使控制疾病基率,不同群体的预测表现仍存在显著差异
- 为医疗AI提示设计提供公平性指导,适合关注AI公平性的研究者
视觉语言模型(VLMs)在医学诊断中展现潜力,但其在使用上下文学习(ICL)时对不同人口统计学子群体的表现尚不明确。本文研究了演示样本的人口统计构成如何影响VLM在两项医学影像任务中的表现:皮肤病变恶性程度预测和胸片中气胸检测。分析显示,ICL通过两种机制影响模型预测:(1) VLM能从提示中学习到子群体特定的疾病基率;(2) 即使控制了子群体疾病基率,模型在不同群体间的预测表现仍存在差异。实证结果为当前VLM的提示设计提供了最佳实践建议(如关注子群体表现、在整体和子群体层面匹配标签基率),同时也指出了提升对这些模型理论理解的下一步方向。
原文摘要 · Abstract (English)
Vision language models (VLMs) show promise in medical diagnosis, but their performance across demographic subgroups when using in-context learning (ICL) remains poorly understood. We examine how the demographic composition of demonstration examples affects VLM performance in two medical imaging tasks: skin lesion malignancy prediction and pneumothorax detection from chest radiographs. Our analysis reveals that ICL influences model predictions through multiple mechanisms: (1) ICL allows VLMs to learn subgroup-specific disease base rates from prompts and (2) ICL leads VLMs to make predictions that perform differently across demographic groups, even after controlling for subgroup-specific disease base rates. Our empirical results inform best-practices for prompting current VLMs (specifically examining demographic subgroup performance, and matching base rates of labels to target distribution at a bulk level and within subgroups), while also suggesting next steps for improving our theoretical understanding of these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。