arXiv:2608.21868cs.AI2026-08

用分层多智能体系统提升临床访谈中抑郁评估的可解释性

HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews

论文配图:HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews
图 1 · 摘自论文原文
  • 分三层智能体协同处理:线索识别、症状评分、整体审核
  • 在E-DAIC数据集上超越现有方法,实现更高准确率与可审计性
  • 适合需要透明决策过程的医疗AI场景,如精神健康筛查

从多模态临床访谈中进行抑郁症评估,需将分散的症状证据整合为连贯的PHQ-8量表。这一过程具有层次性:相关证据常在问答片段中稀疏且依赖上下文,多个问答交互共同支持症状判断,最终评估取决于完整症状谱的连贯性。现有大模型系统或整体处理访谈,或分配通用角色任务,均缺乏显式协调机制来管理证据获取、评分权限、有限反馈与状态记录。为此,我们提出HiMA-MDD,一种对齐评估层级的分层多智能体框架。非代理预处理构建保持上下文的多模态问答单元后,第一层识别候选问答到量表项的关系并支持受限的证据路由;第二层将症状组分配给专业分析师,每名分析师负责一项初步评分;第三层审核完整初步量表,最多请求一轮定向修订,并重构验证后的PHQ-8量表。该设计自然生成分层证据追踪,保留所有中间证据、判断与修订以供审计。最终项目得分确定性生成总分与筛查决策。基于Qwen2.5-72B-Instruct作为核心,实验在E-DAIC数据集上表明,HiMA-MDD优于现有最先进方法。

原文摘要 · Abstract (English)

Depression assessment from multimodal clinical interviews requires integrating dispersed evidence from multiple symptoms into a coherent PHQ-8 profile. This process is hierarchical: relevant evidence is often sparse and context-dependent within local question-answer exchanges, multiple exchanges jointly support symptom-level judgments, and the final assessment depends on the coherence of the complete symptom profile. Existing LLM systems either process interviews holistically or distribute work across generic agent roles; neither design necessarily provides an explicit orchestration mechanism that coordinates evidence access, item-score authority, bounded feedback, and state recording across these levels. To address this gap, we introduce HiMA-MDD, a hierarchical multi-agent harness that aligns this assessment hierarchy with three agent layers. After non-agentic preprocessing constructs context-preserving multimodal QA units, Layer 1 identifies candidate QA-to-item relations and supports bounded item-grounded evidence routing. Layer 2 assigns symptom groups to operational factor specialists, with one specialist responsible for each provisional item score. Layer 3 audits the complete provisional profile, requests at most one round of targeted revision, and reconstructs the verified PHQ-8 profile. This layered design naturally yields a Hierarchical Evidence Trace, preserves all intermediate evidence, judgments, and revisions for auditability. The final item scores then deterministically produce the total score and screening decision. Using Qwen2.5-72B-Instruct as the harness backbone, our experiments on E-DAIC demonstrate that HiMA-MDD outperforms the compared state-of-the-art methods.

抑郁症检测多模态分析可解释AI智能体系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。