通过可验证推理链提升大模型在医疗任务中的可靠性
Haibu Mathematical-Medical Intelligent Agent:Enhancing Large Language Model Reliability in Medical Tasks via Verifiable Reasoning Chains
- 将复杂医疗任务拆解为可验证的原子步骤,实现逻辑与证据可追溯
- 错误检测率超98%,误报率低于1%,显著优于普通大模型
- 支持知识复用与低成本推理,适合高要求医疗场景应用
大型语言模型在医学领域展现出巨大潜力,但容易出现事实性与逻辑性错误,这在高风险医疗环境中不可接受。为此,我们提出「海步数学-医学智能代理」(MMIA),一种基于大模型的架构,通过形式化可验证的推理过程确保可靠性。MMIA将复杂医疗任务递归分解为原子化、基于证据的步骤,整个推理链自动审计逻辑连贯性与证据可追溯性,类似定理证明。核心创新在于「自举模式」,将已验证的推理链存为「定理」,后续任务可通过检索增强生成(RAG)高效求解,从高成本的原始推理转为低开销的验证模式。我们在四个医疗管理领域(包括DRG/DIP审计与医保理赔)进行验证,使用专家标注基准测试。结果表明,MMIA错误检测率超过98%,假阳性率低于1%,显著优于基线大模型。此外,随着知识库成熟,RAG匹配模式预计可使平均处理成本降低约85%。结论表明,MMIA的可验证推理框架是构建可信、透明、低成本人工智能系统的重要一步,使大模型技术在关键医疗应用中具备可行性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) show promise in medicine but are prone to factual and logical errors, which is unacceptable in this high-stakes field. To address this, we introduce the "Haibu Mathematical-Medical Intelligent Agent" (MMIA), an LLM-driven architecture that ensures reliability through a formally verifiable reasoning process. MMIA recursively breaks down complex medical tasks into atomic, evidence-based steps. This entire reasoning chain is then automatically audited for logical coherence and evidence traceability, similar to theorem proving. A key innovation is MMIA's "bootstrapping" mode, which stores validated reasoning chains as "theorems." Subsequent tasks can then be efficiently solved using Retrieval-Augmented Generation (RAG), shifting from costly first-principles reasoning to a low-cost verification model. We validated MMIA across four healthcare administration domains, including DRG/DIP audits and medical insurance adjudication, using expert-validated benchmarks. Results showed MMIA achieved an error detection rate exceeding 98% with a false positive rate below 1%, significantly outperforming baseline LLMs. Furthermore, the RAG matching mode is projected to reduce average processing costs by approximately 85% as the knowledge base matures. In conclusion, MMIA's verifiable reasoning framework is a significant step toward creating trustworthy, transparent, and cost-effective AI systems, making LLM technology viable for critical applications in medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。