arXiv:2504.03135cs.CVcs.AI2025-04被引 11

提出分层提示与解码机制,提升医学图像问答的细粒度理解能力。

Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion

  • 用分层提示对齐文本与图像特征,引导模型聚焦特定区域。
  • 在Rad-Restruct数据集上优于现有方法,层级预测准确率显著提升。
  • 适合需要细粒度医学图像分析的研究者和临床辅助系统开发者。

医学视觉问答(Med-VQA)通过医学图像回答临床问题,有助于诊断决策。构建高效的Med-VQA系统对提升诊断准确性具有重要意义。在此基础上,分层医学VQA通过将问题组织为层次结构,并对不同层级进行专门预测,以处理细粒度差异。尽管已有研究提出分层任务并建立数据集,但仍存在两大挑战:(1) 层次建模不完善导致层级间语义混淆,出现跨层级语义碎片化;(2) 基于Transformer的跨模态自注意力融合方法过度依赖隐式学习,忽视了医疗场景中的关键局部语义关联。为此,本文提出HiCA-VQA框架,包含两个模块:分层提示模块用于预对齐分层文本提示与图像特征,引导模型根据问题类型关注特定图像区域;分层答案解码器则对不同层级问题分别预测,提升多粒度精度。框架还引入交叉注意力融合模块,以图像为查询,文本为键值对。在Rad-Restruct基准上的实验表明,该方法在回答分层细粒度问题上优于现有最先进方法,为分层视觉问答系统提供了有效路径,推动医学图像理解发展。

原文摘要 · Abstract (English)

Medical Visual Question Answering (Med-VQA) answers clinical questions using medical images, aiding diagnosis. Designing the MedVQA system holds profound importance in assisting clinical diagnosis and enhancing diagnostic accuracy. Building upon this foundation, Hierarchical Medical VQA extends Medical VQA by organizing medical questions into a hierarchical structure and making level-specific predictions to handle fine-grained distinctions. Recently, many studies have proposed hierarchical MedVQA tasks and established datasets, However, several issues still remain: (1) imperfect hierarchical modeling leads to poor differentiation between question levels causing semantic fragmentation across hierarchies. (2) Excessive reliance on implicit learning in Transformer-based cross-modal self-attention fusion methods, which obscures crucial local semantic correlations in medical scenarios. To address these issues, this study proposes a HiCA-VQA method, including two modules: Hierarchical Prompting for fine-grained medical questions and Hierarchical Answer Decoders. The hierarchical prompting module pre-aligns hierarchical text prompts with image features to guide the model in focusing on specific image regions according to question types, while the hierarchical decoder performs separate predictions for questions at different levels to improve accuracy across granularities. The framework also incorporates a cross-attention fusion module where images serve as queries and text as key-value pairs. Experiments on the Rad-Restruct benchmark demonstrate that the HiCA-VQA framework better outperforms existing state-of-the-art methods in answering hierarchical fine-grained questions. This study provides an effective pathway for hierarchical visual question answering systems, advancing medical image understanding.

医学图像视觉问答分层建模交叉注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。