arXiv:2511.12446cs.CV2025-11被引 1

让医学视觉问答更可信,测试时自动找图像证据

CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training

  • 用视觉思维链定位关键区域,动态优化答案
  • 不需标注,仅改少量提示词,准确率提升12.3%
  • 适合临床部署,适配多种模型且无需重新训练

医学视觉问答可辅助临床决策,但现有系统在领域迁移下表现不佳,且答案与图像证据关联弱。问题源于模型关注无关区域,且部署时重训练或添加标签不现实。我们提出 CoTBox-TTT,一种基于证据的测试时训练方法,在推理阶段适应视觉语言模型,保持所有主干冻结。该方法仅更新少量连续软提示,通过视觉思维链信号识别与问题相关的图像区域,并增强原图与局部裁剪图上答案的一致性。整个过程无需标签,可插拔适配多种主干模型。在医学 VQA 数据集上的实验表明,该方法适用于实际部署。例如,将 CoTBox-TTT 加入 LLaVA 后,pathVQA 上闭式问答准确率提升 12.3%。

原文摘要 · Abstract (English)

Medical visual question answering could support clinical decision making, yet current systems often fail under domain shift and produce answers that are weakly grounded in image evidence. This reliability gap arises when models attend to spurious regions and when retraining or additional labels are impractical at deployment time. We address this setting with CoTBox-TTT, an evidence-first test-time training approach that adapts a vision-language model at inference while keeping all backbones frozen. The method updates only a small set of continuous soft prompts. It identifies question-relevant regions through a visual chain-of-thought signal and encourages answer consistency across the original image and a localized crop. The procedure is label free, and plug and play with diverse backbones. Experiments on medical VQA show that the approach is practical for real deployments. For instance, adding CoTBox-TTT to LLaVA increases closed-ended accuracy by 12.3% on pathVQA.

医学视觉问答测试时训练视觉链式思维零样本适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。