用多智能体系统提升放射科图像问答的推理准确性与可信度。
A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering
- 设计专用智能体分工处理理解、跨模态推理和答案验证。
- 在模型分歧筛选的难题数据集上超越主流大模型基线。
- 适合需要可解释性与高可靠性的临床AI应用场景。
放射科视觉问答(RVQA)能为胸片图像提供精准回答,减轻放射科医生负担。尽管基于多模态大语言模型(MLLMs)和检索增强生成(RAG)的方法在RVQA中取得进展,但仍面临事实准确性不足、幻觉和跨模态错位等问题。本文提出一种多智能体系统(MAS),支持复杂推理,包含专门负责上下文理解、多模态推理和答案验证的智能体。我们在通过模型分歧过滤构建的挑战性RVQA数据集上评估该系统,该数据集包含多个MLLM一致难以解答的困难案例。大量实验表明,该系统显著优于强基线模型;案例研究进一步展示了其可靠性与可解释性。本工作凸显了多智能体方法在支持需复杂推理的可解释、可信临床AI应用中的潜力。
原文摘要 · Abstract (English)
Radiology visual question answering (RVQA) provides precise answers to questions about chest X-ray images, alleviating radiologists' workload. While recent methods based on multimodal large language models (MLLMs) and retrieval-augmented generation (RAG) have shown promising progress in RVQA, they still face challenges in factual accuracy, hallucinations, and cross-modal misalignment. We introduce a multi-agent system (MAS) designed to support complex reasoning in RVQA, with specialized agents for context understanding, multimodal reasoning, and answer validation. We evaluate our system on a challenging RVQA set curated via model disagreement filtering, comprising consistently hard cases across multiple MLLMs. Extensive experiments demonstrate the superiority and effectiveness of our system over strong MLLM baselines, with a case study illustrating its reliability and interpretability. This work highlights the potential of multi-agent approaches to support explainable and trustworthy clinical AI applications that require complex reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。