通过视觉概念排序发现医学大模型的潜在偏见与捷径学习行为
Visual concept ranking uncovers medical shortcuts used by large multimodal models
- 提出视觉概念排序方法,识别模型依赖的关键视觉特征
- 发现模型在不同人群间表现差异显著,存在诊断偏差
- 适用于医疗AI审计,帮助研究人员发现模型决策捷径
在医疗等高风险领域,确保机器学习模型的可靠性需依赖可审计的方法以揭示其缺陷。本文提出一种用于识别大型多模态模型(LMMs)中重要视觉概念的方法,并应用于医学任务中的模型行为分析。主要聚焦于临床皮肤科图像中恶性皮肤病变的分类任务,辅以胸部X光片和自然图像的补充实验。当模型使用示例演示进行提示时,发现其在不同人口统计子群体间表现存在意外差距。应用所提出的视觉概念排名(VCR)方法后,生成关于不同视觉特征依赖关系的假设,并通过人工干预验证其有效性。
原文摘要 · Abstract (English)
Ensuring the reliability of machine learning models in safety-critical domains such as healthcare requires auditing methods that can uncover model shortcomings. We introduce a method for identifying important visual concepts within large multimodal models (LMMs) and use it to investigate the behaviors these models exhibit when prompted with medical tasks. We primarily focus on the task of classifying malignant skin lesions from clinical dermatology images, with supplemental experiments including both chest radiographs and natural images. After showing how LMMs display unexpected gaps in performance between different demographic subgroups when prompted with demonstrating examples, we apply our method, Visual Concept Ranking (VCR), to these models and prompts. VCR generates hypotheses related to different visual feature dependencies, which we are then able to validate with manual interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。