用因果图提升医疗文献筛选准确率,零幻觉且可解释。
Causal-Enhanced AI Agents for Medical Research Screening
- 构建双层知识图谱,强制每条因果结论来自文献证据。
- 在阿尔茨海默病运动研究中实现95%准确率和0%幻觉。
- 适合需要高可信度的医疗AI应用,尤其关注可解释性。
系统性综述是循证医学的关键,但每年超150万篇论文手动筛选不可行。当前AI在综述任务中存在幻觉问题,早期模型幻觉率28%-40%,现代模型降至2%-15%,在影响患者安全时仍不可接受。本文提出因果图增强的检索增强生成系统,结合显式因果推理与双层知识图谱,强制执行证据优先原则,所有因果主张均追溯至检索到的文献,并自动生成有向无环图展示干预-结果路径。在234篇阿尔茨海默病运动研究摘要上的评估显示,CausalAgent达到95%准确率、100%召回率、0%幻觉,远优于基线模型(34%准确率,10%幻觉)。自动生成的因果图支持机制建模、可视化合成与可解释性提升。尽管本验证聚焦于10个关于痴呆运动研究的问题,其架构设计展示了可信医疗AI的可迁移原则,凸显因果推理在高风险医疗场景中的潜力。
原文摘要 · Abstract (English)
Systematic reviews are essential for evidence-based medicine, but reviewing 1.5 million+ annual publications manually is infeasible. Current AI approaches suffer from hallucinations in systematic review tasks, with studies reporting rates ranging from 28--40% for earlier models to 2--15% for modern implementations which is unacceptable when errors impact patient care. We present a causal graph-enhanced retrieval-augmented generation system integrating explicit causal reasoning with dual-level knowledge graphs. Our approach enforces evidence-first protocols where every causal claim traces to retrieved literature and automatically generates directed acyclic graphs visualizing intervention-outcome pathways. Evaluation on 234 dementia exercise abstracts shows CausalAgent achieves 95% accuracy, 100% retrieval success, and zero hallucinations versus 34% accuracy and 10% hallucinations for baseline AI. Automatic causal graphs enable explicit mechanism modeling, visual synthesis, and enhanced interpretability. While this proof-of-concept evaluation used ten questions focused on dementia exercise research, the architectural approach demonstrates transferable principles for trustworthy medical AI and causal reasoning's potential for high-stakes healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。