arXiv:2511.01458cs.CVcs.AI2025-11被引 2

提出新方法提升手术视觉问答的可靠性,让系统更懂问题才敢给答案。

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA

  • 通过双向门控融合问题与答案语义匹配度,改进不确定性评估
  • 在同图不同问下,对错误回答的识别率提升最高达21%
  • 适合临床部署的视觉问答系统,尤其关注语言变化的场景

手术中视觉问答系统的安全与可靠至关重要,错误或模糊的回答可能导致患者伤害。现有不确定性估计方法(如语义最近邻熵,SNNE)未显式考虑问题条件,可能对语义一致但与临床问题不符的答案赋予过高置信度,尤其在问题表述变化时。本文提出问题对齐的语义最近邻熵(QA-SNNE),一种黑盒不确定性估计器,通过双边门控将问题-答案对齐纳入语义熵计算,基于嵌入、蕴含或交叉编码策略加权答案间的语义相似性。为测试语言变化鲁棒性,构建了基准手术VQA数据集的非模板重述版本,仅修改问题措辞而保持图像和真实答案不变。在两个基准手术VQA数据集上,对五种VQA模型在零样本和参数高效微调(PEFT)设置下进行评估,包括非模板问题。QA-SNNE在端点可视化18-VQA数据集上,使三种零样本模型中的两种在模板内问题上的AUROC提升达15%(如Llama3.2)至21%(如Qwen2.5),在非模板重述下最高提升8%,外部验证结果混合。总体而言,QA-SNNE提供了一种实用、模型无关的安全保障,将语义不确定性与问题相关性紧密关联。

原文摘要 · Abstract (English)

Safety and reliability are critical for deploying visual question answering (VQA) systems in surgery, where incorrect or ambiguous responses can cause patient harm. A key limitation of existing uncertainty estimation methods, such as Semantic Nearest Neighbor Entropy (SNNE), is that they do not explicitly account for the conditioning question. As a result, they may assign high confidence to answers that are semantically consistent yet misaligned with the clinical question, especially under variation in question phrasing. We propose Question-Aligned Semantic Nearest Neighbor Entropy (QA-SNNE), a black-box uncertainty estimator that incorporates question-answer alignment into semantic entropy through bilateral gating. QA-SNNE measures uncertainty by weighting pairwise semantic similarities among sampled answers according to their relevance to the question, using embedding-based, entailment-based, or cross-encoder alignment strategies. To assess robustness to language variation, we construct an out-of-template rephrased version of a benchmark surgical VQA dataset, where only the question wording is modified while images and ground-truth answers remain unchanged. We evaluate QA-SNNE on five VQA models across two benchmark surgical VQA datasets in both zero-shot and parameter-efficient fine-tuned (PEFT) settings, including out-of-template questions. QA-SNNE improves AUROC on EndoVis18-VQA for two of three zero-shot models in-template (e.g., +15% for Llama3.2 and +21% for Qwen2.5) and achieves up to +8% AUROC improvement under out-of-template rephrasing, with mixed results on external validation. Overall, QA-SNNE provides a practical, model-agnostic safeguard for surgical VQA by linking semantic uncertainty to question relevance.

视觉问答医疗AI不确定性估计手术辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。