轻量级医学问答模型,用低频特征+图注意力提升效率与准确率
MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

- 用DCT低频特征结合可学习记忆库增强视觉表征
- 在9种影像模态上超200万合成数据集上表现超越大模型
- 适合资源受限的临床场景部署,计算成本显著降低
医学视觉问答(Med-VQA)在临床决策支持中潜力巨大,但受限于标注数据少和现有大模型计算开销高。本文提出MedFG-VQA,一种轻量级框架,通过记忆库增强基于DCT的低频特征,并采用图增强交叉注意力实现有效的视觉-文本对齐。核心包括:频率-记忆融合(FMF),通过从基于DCT分解构建的可学习记忆库中检索来增强低频特征;图感知交叉注意力(GACA),利用交叉注意力对齐视觉-文本特征,并通过图卷积聚合进行优化。为缓解数据稀缺问题,构建了包含超过200万组问答对的SynMed-VQA合成数据集,覆盖9种影像模态和10个主要器官,由GPT-4o生成。在SynMed-VQA及三个标准生物医学VQA基准上的实验表明,MedFG-VQA在性能上可媲美甚至优于更大模型,同时计算成本显著降低,展现出良好的临床部署潜力。
原文摘要 · Abstract (English)
Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language models. We propose MedFG-VQA, a lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment. Specifically, our approach features two key components: Frequency-Memory Fusion (FMF), which enhances low-frequency features by retrieving from a learnable memory bank built on DCT decomposition, and Graph-Aware Cross-Attention (GACA), which aligns visual-textual features via cross-attention and refines them through graph-convolutional aggregation. To address data scarcity, we construct SynMed-VQA, a large-scale synthetic dataset comprising over 2 million question-answer pairs across 9 imaging modalities and 10 major organs, generated with GPT-4o. Extensive experiments on SynMed-VQA and three other standard biomedical VQA benchmarks demonstrate that MedFG-VQA achieves competitive or superior performance compared to much larger models while maintaining significantly lower computational costs, highlighting its efficiency and potential for clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。