arXiv:2605.01911cs.CV2026-05被引 1

揭露手术视觉问答中模型依赖语言线索而非真看图的真相

SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?

论文配图:SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?
图 1 · 摘自论文原文
  • 设计配对问题对比,剥离实体名称检验模型是否真看图
  • 五种模型在去偏问题上性能普遍下降,证明严重依赖语言捷径
  • 适合关注医疗AI可解释性与评测公平性的研究者参考

视觉语言模型(VLMs)在手术视觉问答(VQA)中表现优异,但现有数据集常存在语言捷径——问题表述隐含答案范围。这使得模型性能可能源于语言线索而非真实视觉理解。为此,我们提出SurgCheck,一个诊断性基准,通过配对问题设计:同一手术帧对应包含实体名的原始问题和移除实体名但保留相同视觉内容与正确答案的去偏版本。性能差距反映模型对语言捷径的依赖程度。为确保去偏问题仍具可答性,引入四种定位提示:边界框、箭头、空间位置和迂回表达。我们在零样本和微调设置下评估通用及专用VLMs。针对开放式零样本回答,引入大模型作为裁判的评估协议。结果表明,五种VLM在去偏问题上均出现显著性能下降,而仅用文本的消融实验显示动作和目标预测性能几乎不变,说明其主要依赖语言捷径而非视觉推理。结论:SurgCheck提供了一个可控的诊断框架,揭示了现有基准中被语言偏差掩盖的失败模式。强性能不等于真理解,强调需在手术VQA中引入偏差意识的评估。

原文摘要 · Abstract (English)

Purpose: Vision-language models (VLMs) have shown promising performance in surgical visual question answering (VQA). However, existing surgical VQA datasets often contain linguistic shortcuts, where question phrasing implicitly constrains the answer space. It remains unclear whether reported performance reflects visual understanding or reliance on such linguistic shortcuts. Methods: We introduce SurgCheck, a diagnostic benchmark for quantifying linguistic shortcut reliance in surgical VQA. SurgCheck employs a paired-question design in which each surgical frame is associated with an original question containing entity names and a less-biased counterpart that removes these names while preserving identical visual content and ground-truth answers. The resulting performance gap provides a diagnostic signal of shortcut reliance. To ensure that the less-biased question remains well-defined even without entity names, four grounding cues are incorporated: bounding box, arrow, spatial position, and periphrasis. We evaluate both general-purpose and surgical-specific VLMs under zero-shot and fine-tuned settings on SurgCheck. To evaluate open-ended zero-shot responses, we introduce an LLM-as-a-judge evaluation protocol. Results: Using SurgCheck, we observe consistent performance degradation on less-biased questions across five VLMs, despite identical visual inputs. Text-only ablation reveals minimal performance drops for action and target prediction, indicating that action and target prediction is largely driven by linguistic shortcuts rather than visual reasoning. Conclusion: SurgCheck provides a controlled diagnostic framework that exposes failure modes masked by linguistic bias in existing surgical VQA benchmarks. Our findings demonstrate that strong benchmark performance does not necessarily imply faithful visual understanding, underscoring the need for bias-aware evaluation in surgical VQA.

视觉问答医疗AI模型评估语言捷径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。