医学视觉语言模型在真实与从众间存在权衡,难以同时避免幻觉和盲从。
To Agree or To Be Right? The Grounding-Sycophancy Tradeoff in Medical Vision-Language Models
- 提出三类新指标,量化模型的接地性与顺从性
- 7-8B参数模型在1151个测试中无一达到安全指数0.35以上
- 医疗专用模型更抗压力但幻觉更多,通用模型最顺从
经过医学领域适配的视觉语言模型(VLMs)在视觉问答任务上表现优异,但其对幻觉和顺从性这两种关键失败模式的鲁棒性仍不明确,尤其在两者共现时。我们在三个医学VQA数据集上评估了六种VLMs(三种通用型,三种医疗专用型),发现存在接地性与顺从性之间的权衡:幻觉最少的模型最易顺从,而最具抗压能力的模型幻觉率高于所有医疗专用模型。为刻画该权衡,我们提出三个指标:L-VASE(VASE的对数空间重构,避免双重归一化)、CCS(置信度校准的顺从性评分,惩罚高自信屈服)、临床安全指数(CSI),通过几何平均融合接地性、自主性和校准性。在1151个测试案例中,没有任何模型的CSI超过0.35,表明这些7-8B参数的VLM均无法同时具备强接地性与抗社会压力能力。研究建议在临床应用前必须联合评估这两项属性。代码已开源。
原文摘要 · Abstract (English)
Vision-language models (VLMs) adapted to the medical domain have shown strong performance on visual question answering benchmarks, yet their robustness against two critical failure modes, hallucination and sycophancy, remains poorly understood, particularly in combination. We evaluate six VLMs (three general-purpose, three medical-specialist) on three medical VQA datasets and uncover a grounding-sycophancy tradeoff: models with the lowest hallucination propensity are the most sycophantic, while the most pressure-resistant model hallucinates more than all medical-specialist models. To characterize this tradeoff, we propose three metrics: L-VASE, a logit-space reformulation of VASE that avoids its double-normalization; CCS, a confidence-calibrated sycophancy score that penalizes high-confidence capitulation; and Clinical Safety Index (CSI), a unified safety index that combines grounding, autonomy, and calibration via a geometric mean. Across 1,151 test cases, no model achieves a CSI above 0.35, indicating that none of the evaluated 7-8B parameter VLMs is simultaneously well-grounded and robust to social pressure. Our findings suggest that joint evaluation of both properties is necessary before these models can be considered for clinical use. Code is available at https://github.com/UTSA-VIRLab/AgreeOrRight
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。