arXiv:2504.11777cs.CVcs.LG2025-04

用大模型生成语义一致的问答变体,提升医疗视觉问答一致性。

Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets

  • 用大模型生成语义等价的多样化问题,增强数据多样性。
  • 增强数据后,模型准确率平均提升19.35%,一致性指标提升11.61%。
  • 适合医疗AI研究者和需要提升模型鲁棒性的开发者。

医学视觉问答(MVQA)系统能根据自然语言问题解读医学图像,但问题表述的语言差异常导致结果不一致。为此,本文提出语义等价问题增强(SEQA)框架,利用大语言模型生成多样且语义一致的问题重述,以提升语言多样性同时保持语义不变。进一步提出总一致率-语义等价输入与正确答案(TAR-SC)评估指标,衡量模型对语义等价变体的响应一致性。此外还引入三个多样性度量:每张图像的平均问答对数(ANQI)、相同答案的问题平均数(ANQA)和相同语义的开放问题平均数(ANQS)。基于SEQA框架,对SLAKE、VQA-RAD和PathVQA三个公开数据集进行增强,结果显示:平均ANQI提升86.1,ANQA提升85.1,ANQS提升46。在增强数据上对M2I2、MUMC和BiomedGPT三个模型进行零样本与微调实验,微调模型平均准确率提升19.35%,TAR-SC指标平均提升11.61%,表明模型一致性显著增强。

原文摘要 · Abstract (English)

Medical Visual Question Answering (MVQA) systems can interpret medical images in response to natural language queries. However, linguistic variability in question phrasing often undermines the consistency of these systems. To address this challenge, we propose a Semantically Equivalent Question Augmentation (SEQA) framework, which leverages large language models (LLMs) to generate diverse yet semantically equivalent rephrasings of questions. Specifically, this approach enriches linguistic diversity while preserving semantic meaning. We further introduce an evaluation metric, Total Agreement Rate with Semantically Equivalent Input and Correct Answer (TAR-SC), which assesses a model's capability to generate consistent and correct responses to semantically equivalent linguistic variations. In addition, we also propose three other diversity metrics - average number of QA items per image (ANQI), average number of questions per image with the same answer (ANQA), and average number of open-ended questions per image with the same semantics (ANQS). Using the SEQA framework, we augmented the benchmarked MVQA public datasets of SLAKE, VQA-RAD, and PathVQA. As a result, all three datasets achieved significant improvements by incorporating more semantically equivalent questions: ANQI increased by an average of 86.1, ANQA by 85.1, and ANQS by 46. Subsequent experiments evaluate three MVQA models (M2I2, MUMC, and BiomedGPT) under both zero-shot and fine-tuning settings on the enhanced datasets. Experimental results in MVQA datasets show that fine-tuned models achieve an average accuracy improvement of 19.35%, while our proposed TAR-SC metric shows an average improvement of 11. 61%, indicating a substantial enhancement in model consistency.

医疗AI视觉问答大模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。