用频谱分析与量子检索增强医学图文问答,提升准确率与可解释性。
Q-FSRU: Quantum-Augmented Frequency-Spectral Fusion for Medical Visual Question Answering
- 将图像与文本特征转至频域,通过快速傅里叶变换过滤噪声。
- 引入量子启发检索机制,从外部知识源获取真实医学事实。
- 在VQA-RAD数据集上超越旧模型,尤其擅长复杂跨模态推理。
解决需同时理解图像与文本的临床难题仍是医疗AI的重大挑战。本文提出Q-FSRU,结合频谱表示与融合(FSRU)及量子检索增强生成(Quantum RAG)方法,用于医学视觉问答(VQA)。模型接收医学图像与相关文本特征,经快速傅里叶变换(FFT)转入频域,聚焦有意义信息并抑制噪声。为提升准确性并确保答案基于真实知识,引入量子启发检索系统,利用量子相似性技术从外部来源获取医学事实。这些知识与频域特征融合,增强推理能力。在包含真实放射科图像与问题的VQA-RAD数据集上评估,结果表明Q-FSRU优于先前模型,尤其在需要图像-文本联合推理的复杂病例中表现突出。频谱与量子信息的结合提升了性能与可解释性。该方法为构建智能、清晰、有益于医生的AI工具提供了有前景的路径。
原文摘要 · Abstract (English)
Solving tough clinical questions that require both image and text understanding is still a major challenge in healthcare AI. In this work, we propose Q-FSRU, a new model that combines Frequency Spectrum Representation and Fusion (FSRU) with a method called Quantum Retrieval-Augmented Generation (Quantum RAG) for medical Visual Question Answering (VQA). The model takes in features from medical images and related text, then shifts them into the frequency domain using Fast Fourier Transform (FFT). This helps it focus on more meaningful data and filter out noise or less useful information. To improve accuracy and ensure that answers are based on real knowledge, we add a quantum-inspired retrieval system. It fetches useful medical facts from external sources using quantum-based similarity techniques. These details are then merged with the frequency-based features for stronger reasoning. We evaluated our model using the VQA-RAD dataset, which includes real radiology images and questions. The results showed that Q-FSRU outperforms earlier models, especially on complex cases needing image-text reasoning. The mix of frequency and quantum information improves both performance and explainability. Overall, this approach offers a promising way to build smart, clear, and helpful AI tools for doctors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。