arXiv:2608.28624cs.CLcs.AI2026-08

用多智能体增强生成,让AI更准确地总结帕金森病长期临床数据。

MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson's Disease Assessments

论文配图:MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson's Disease Assessments
图 1 · 摘自论文原文
  • 拆分任务为领域专用智能体,结合检索与验证生成临床结论。
  • 事实精确率提升122%,幻觉率降低98%,时间一致性显著改善。
  • 适合需要精准医疗摘要的医生和研究者使用。

帕金森病单次及纵向临床评估的准确解读耗时且依赖专科知识。尽管大语言模型(LLMs)可生成自然语言摘要,但常缺乏特定临床背景,难以对结构化纵向数据生成事实正确且时间连贯的回应。为此,我们提出MA-RAG,一种查询驱动的多智能体检索增强生成框架,将临床推理分解为领域专用智能体,结合结构化事实提取,并通过最终验证阶段合成具有临床依据的摘要。该框架支持四种临床分析任务:单次会诊、轨迹、对比与队列总结。我们采用客观指标(事实精确率、幻觉率、时间保真度、语义相似度)及临床专家主观评价进行评估。相比传统方法、仅使用RAG及单智能体RAG基线,MA-RAG在事实正确性上大幅提升:事实精确率从0.436增至0.990(相对提升122%),幻觉率从0.564降至0.010(降低98%),同时在组织性和临床实用性方面获得专家最高评分。结果表明,领域专用多智能体推理能实现结构化纵向临床数据的可靠查询驱动摘要。

原文摘要 · Abstract (English)

Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson's disease is time-consuming and often depends on specialist expertise. Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-specific clinical grounding and struggle to produce factually correct and temporally consistent responses for structured longitudinal assessment data. To address these limitations, we propose MA-RAG, a query-driven multi-agent retrieval-augmented generation framework that decomposes clinical reasoning into domain-specialized agents, combines structured fact extraction, and synthesizes clinically grounded summaries through a final verification stage. The framework supports four clinical analysis tasks: single-session, trajectory, comparison, and cohort summarization. We evaluate MA-RAG using objective metrics, namely Fact Precision, Hallucination Rate, Temporal Fidelity, and Semantic Similarity, together with subjective evaluations conducted by clinical experts. Compared to Traditional, RAG-only, and Single-agent RAG baselines, MA-RAG substantially improves factual correctness, achieving up to a 122% relative increase in Fact Precision (from 0.436 to 0.990) and reducing the Hallucination Rate by up to 98% (from 0.564 to 0.010), while consistently receiving top ratings from clinical experts for organization and clinical usefulness. These results demonstrate that domain-specialized multi-agent reasoning enables reliable query-driven summarization of structured longitudinal clinical assessment data.

医疗AI多智能体临床摘要生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。