通过频域双分支融合,提升医学图像问答的细粒度对齐能力。
Frequency-Domain Dual-Branch Fusion for Medical Visual Question Answering

- 在频域中分两支处理视觉与语言特征,根据问题自适应选择高低频信息
- 在VQA-RAD和SLAKE上准确率提升3.2%和4.1%,且模型轻量高效
- 适合需要精准医学图像理解的临床辅助系统开发者
医学视觉问答(VQA)需对齐细微的视觉证据,如病灶纹理、边界锐度及弥漫性密度变化与临床语言。现有空间域多模态融合方法难以充分挖掘视觉与文本表示中的互补频域信息。本文提出一种双分支频域融合模块,根据输入问题动态调节频谱过滤,先在频域中自适应选择全局低频结构与细粒度高频细节,再重建空间表示以生成答案。为提供更丰富的频谱支持,从冻结的BiomedCLIP编码器的早期纹理敏感层与最终语义层提取互补特征,并使用对称InfoNCE目标对齐两者与问题表示,随后与BioBART解码器进行分阶段联合训练。模型在PMC-VQA上预训练,在VQA-RAD与SLAKE基准上微调,结果表明,具备频域感知的多模态融合显著提升医学VQA性能,同时保持轻量高效的架构。
原文摘要 · Abstract (English)
Medical Visual Question Answering (VQA) requires aligning subtle visual evidence, including lesion texture, boundary sharpness, and diffuse density changes, with clinical language. Existing multimodal fusion approaches operating in the spatial domain may not fully exploit complementary frequency information present in visual and textual representations. We introduce a dual-branch frequency-domain fusion module that conditions spectral filtering on the input question, enabling adaptive selection of global low-frequency structure and fine-grained high-frequency detail before reconstructing the spatial representation for answer generation. To provide a richer spectrum for filtering, we extract complementary features from early texture-sensitive and final semantic layers of a frozen BiomedCLIP encoder and align both with the question representation using a symmetric InfoNCE objective prior to staged joint training with a BioBART decoder. We pretrain the proposed model on PMC-VQA and fine-tune it on the VQA-RAD and SLAKE benchmarks, demonstrating that frequency-aware multimodal fusion improves medical VQA performance while maintaining a lightweight and efficient architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。