arXiv:2607.17834cs.CVcs.AI2026-07

提出新基准与方法,提升内镜问答中复杂答案的一致性。

Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA

论文配图:Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA
图 1 · 摘自论文原文
  • 构建配对的复杂-原子答案一致性评测集,评估模型答案一致性。
  • 部分模型复杂答案准确率高但原子答案与一致性表现差。
  • 提出无需训练的修正与择优机制,提升答案可信度。

内镜视觉问答(VQA)日益涉及需整合多个答案成分的复杂问题,而非孤立的事实查询。此类复杂答案可能被判定为正确,即使同一模型在对应原子问题上失败。本文提出EndoCA,一个配对的复杂-原子答案一致性评测基准,用于评估复杂答案是否与同图原子答案保持一致。EndoCA包含两个模块:EndoCA-Core评估实际内镜VQA中常见的紧凑问题复杂度模式;EndoCA-Diagnostic支持在递增复杂度下的受控分析。在11个跨类型VLMs上评估发现,部分模型虽具高复杂答案准确率,但原子答案准确率及复杂-原子一致性显著偏低。为此,提出无需训练的原子支持修正(ASR)机制:利用模型生成的原子答案作为上下文前提进行答案修正,并实现一致性引导的择优回答。在四个公开可用模型上,ASR-Revise在复杂答案准确率小幅变化下提升配对正确率;ASR-Selective通过允许模型放弃不可靠情况,提升已答案例的准确性。EndoCA与ASR共同提供一致性感知的评测与无训练修复机制,推动内镜VQA答案可信度提升。

原文摘要 · Abstract (English)

Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather than isolated factual queries. Such complex answers may be scored as correct even when the same model fails on associated atomic questions. We introduce EndoCA, a paired complex-atomic answer consistency benchmark for evaluating whether complex answers remain consistent with same-image atomic answers. EndoCA contains two suites: EndoCA-Core evaluates compact question-complexity patterns commonly seen in practical endoscopic VQA, and EndoCA-Diagnostic supports controlled analysis across increasing question complexity. We evaluate 11 VLMs spanning open, medical, endoscopy-adapted, and closed-source models on EndoCA. Some VLMs achieve high complex-answer accuracy, yet their atomic-answer accuracy and complex-atomic answer consistency remain substantially lower. To reduce this complex-atomic inconsistency, we introduce Atomic-Support Reconciliation (ASR), a training-free mechanism that uses model-generated atomic answers as contextual premises for answer revision and consistency-guided selective answering. On four selected publicly available models, ASR-Revise improves paired complex-atomic correctness with modest changes in complex-answer accuracy, while ASR-Selective improves accuracy on answered cases by allowing the model to abstain from less reliable cases. Together, EndoCA and ASR provide a consistency-aware benchmark and a training-free mechanism for answer reconciliation and selective answering in endoscopic VQA.

内镜VQA答案一致性视觉语言模型医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。