多智能体框架提升心理筛查准确性与可解释性
A Multi-Agent Audit Framework for High-Stakes Reasoning: Evaluation and Interpretability in Clinical Mental Health Screening

- 分步构建感知、检索生成、推理和审计四阶段协作流程
- 在DAIC-WOZ数据集上将抑郁评分误差降低至5.02
- 适合医疗AI决策支持系统开发者参考
高风险推理任务需要透明可验证的流程,但传统单模型大语言模型在零样本条件下易产生幻觉且可解释性差。为此,我们提出一种多智能体审计框架,模拟多步骤协同验证过程。基于模块化LangChain工作流,在临床心理健康筛查领域进行实证验证。框架将推理过程分解为感知、知识检索增强生成(RAG)、链式思维(CoT)临床推断及关键审计验证四个阶段。在本地部署开源模型并使用DAIC-WOZ数据集评估,结果表明该多智能体管道显著优于单智能体基线,将PHQ-8抑郁严重度预测的均值绝对误差(MAE)从5.35降至5.02。通过暴露跨智能体验证轨迹,框架有效缓解推理偏差,提供高度可解释的诊断依据,为非单一模型扩展的可靠AI辅助决策支持提供了通用范式。数据与代码已开源以保障可复现性。
原文摘要 · Abstract (English)
High-stakes reasoning tasks necessitate transparent and verifiable workflows, yet conventional single-model large language models (LLMs) often struggle with hallucination and low interpretability under zero-shot paradigms. To address this general AI challenge, we propose a Multi-Agent Audit Framework that simulates a collaborative, multi-step verification process. We empirically validate this architecture in the sensitive domain of clinical mental health screening using a modular LangChain workflow. Our framework decomposes the reasoning process into a Perception Agent, Knowledge Retrieval-Augmented Generation (RAG), Chain-of-Thought (CoT) clinical inference, and a critical Audit verification stage. We evaluated this framework on the DAIC-WOZ dataset using locally deployed open-source models. Experimental results demonstrate that our multi-agent pipeline significantly outperforms single-agent baselines, reducing the Mean Absolute Error (MAE) for PHQ-8 depression severity prediction from 5.35 to 5.02. By exposing cross-agent validation traces, the framework mitigates reasoning drift and provides highly interpretable diagnostic rationales, offering a generalizable paradigm for reliable AI-assisted decision support beyond isolated model scaling. We make data and code open access on GitHub for replicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。