arXiv:2503.02476cs.CVcs.AI2025-03被引 1

提出双层次语义一致性框架,提升医学视觉问答的图文对齐效果

BioD2C: A Dual-level Semantic Consistency Constraint Framework for Biomedical VQA

  • 在模型与特征双层实现图文语义对齐,动态调整视觉特征
  • 新数据集BioVGQ过滤人为修改图像,提升问答对与多模态上下文一致性
  • 在多个下游数据集上达到当前最优性能,适合医学AI研究者参考

医学视觉问答(Biomedical VQA)在辅助医疗诊断等领域具有重要应用价值。然而现有模型仅在大语言模型层面进行多模态信息交互,导致复杂任务下语义对齐不足。为此,本文提出BioD2C:一种双层次语义一致性约束框架,在模型与特征两个层面实现多模态交互对齐,使模型能根据问题自适应学习视觉特征。首先通过图像-文本融合机制将文本特征注入视觉特征,实现特征级语义交互;随后引入基于文本队列的跨模态软语义损失函数,进一步对齐图像与问题语义。本文构建新数据集BioVGQ,通过人工修正图像筛选和问答对与多模态上下文对齐,缓解已有数据集的固有偏差。大量实验表明,BioD2C在多个下游数据集上达到当前最优性能,展现出优异的鲁棒性、泛化能力,具备推动医学VQA研究的潜力。

原文摘要 · Abstract (English)

Biomedical visual question answering (VQA) has been widely studied and has demonstrated significant application value and potential in fields such as assistive medical diagnosis. Despite their success, current biomedical VQA models perform multimodal information interaction only at the model level within large language models (LLMs), leading to suboptimal multimodal semantic alignment when dealing with complex tasks. To address this issue, we propose BioD2C: a novel Dual-level Semantic Consistency Constraint Framework for Biomedical VQA, which achieves dual-level semantic interaction alignment at both the model and feature levels, enabling the model to adaptively learn visual features based on the question. Specifically, we firstly integrate textual features into visual features via an image-text fusion mechanism as feature-level semantic interaction, obtaining visual features conditioned on the given text; and then introduce a text-queue-based cross-modal soft semantic loss function to further align the image semantics with the question semantics. Specifically, in this work, we establish a new dataset, BioVGQ, to address inherent biases in prior datasets by filtering manually-altered images and aligning question-answer pairs with multimodal context, and train our model on this dataset. Extensive experimental results demonstrate that BioD2C achieves state-of-the-art (SOTA) performance across multiple downstream datasets, showcasing its robustness, generalizability, and potential to advance biomedical VQA research.

医学视觉问答多模态对齐数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。