用频谱分析与量子检索提升医疗图文问答准确率
Q-FSRU: Quantum-Augmented Frequency-Spectral For Medical Visual Question Answering
- 将医学图像与文本转至频域,抑制噪声增强关键特征
- 结合量子启发检索获取真实医学知识,提升答案可信度
- 适合需要跨模态推理的临床AI应用,解释性更强
解决需同时理解图像与文本的复杂临床问题仍是医疗AI的重大挑战。本文提出Q-FSRU模型,融合频谱表示与融合(FSRU)及量子检索增强生成(Quantum RAG)方法,用于医疗视觉问答(VQA)。该模型接收医学图像与相关文本特征,通过快速傅里叶变换(FFT)将其转换至频域,从而聚焦有意义信息并过滤噪声。为提升准确性并确保答案基于真实知识,引入量子启发检索系统,利用量子相似性技术从外部来源获取医学事实。这些知识与频域特征融合,实现更强推理能力。在包含真实放射科图像与问题的VQA-RAD数据集上评估显示,Q-FSRU在需图像-文本联合推理的复杂案例中优于先前模型。频谱与量子信息的结合提升了性能与可解释性。该方法为构建智能、清晰且对医生有帮助的AI工具提供了新思路。
原文摘要 · Abstract (English)
Solving tough clinical questions that require both image and text understanding is still a major challenge in healthcare AI. In this work, we propose Q-FSRU, a new model that combines Frequency Spectrum Representation and Fusion (FSRU) with a method called Quantum Retrieval-Augmented Generation (Quantum RAG) for medical Visual Question Answering (VQA). The model takes in features from medical images and related text, then shifts them into the frequency domain using Fast Fourier Transform (FFT). This helps it focus on more meaningful data and filter out noise or less useful information. To improve accuracy and ensure that answers are based on real knowledge, we add a quantum inspired retrieval system. It fetches useful medical facts from external sources using quantum-based similarity techniques. These details are then merged with the frequency-based features for stronger reasoning. We evaluated our model using the VQA-RAD dataset, which includes real radiology images and questions. The results showed that Q-FSRU outperforms earlier models, especially on complex cases needing image text reasoning. The mix of frequency and quantum information improves both performance and explainability. Overall, this approach offers a promising way to build smart, clear, and helpful AI tools for doctors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。