arXiv:2501.13951cs.CLcs.AI2025-01被引 4

用多个专家模型分层处理长文本心理评估,减少错误和幻觉。

A Layered Multi-Expert Framework for Long-Context Mental Health Assessments

  • 分层设计:早期任务拆解,后期融合优化,多模型协同决策。
  • 在DAIC-WOZ数据集上,准确率与F1分数均优于单模型基准,PHQ-8误差下降。
  • 适合高风险心理筛查场景,提升AI评估的可信度与临床实用性。

长文本心理健康评估对大语言模型(LLMs)构成独特挑战,常因处理长上下文而产生幻觉或推理不一致。本文提出分层多专家框架SMMR,通过多个LLM与专用小型模型作为平等‘专家’协同工作。早期层将任务分解为短小、离散的子任务,后期层利用更强大的长上下文模型整合并优化部分输出。我们在DAIC-WOZ抑郁症筛查数据集及48个经精神科诊断的案例研究中评估SMMR,结果表明其在准确率、F1分数和PHQ-8误差方面均持续优于单模型基线。通过引入多样化的‘第二意见’,SMMR有效缓解了幻觉,捕捉细微临床特征,提升了高风险心理评估中的可靠性。研究结果凸显多专家框架在可信AI辅助筛查中的价值。

原文摘要 · Abstract (English)

Long-form mental health assessments pose unique challenges for large language models (LLMs), which often exhibit hallucinations or inconsistent reasoning when handling extended, domain-specific contexts. We introduce Stacked Multi-Model Reasoning (SMMR), a layered framework that leverages multiple LLMs and specialized smaller models as coequal 'experts'. Early layers isolate short, discrete subtasks, while later layers integrate and refine these partial outputs through more advanced long-context models. We evaluate SMMR on the DAIC-WOZ depression-screening dataset and 48 curated case studies with psychiatric diagnoses, demonstrating consistent improvements over single-model baselines in terms of accuracy, F1-score, and PHQ-8 error reduction. By harnessing diverse 'second opinions', SMMR mitigates hallucinations, captures subtle clinical nuances, and enhances reliability in high-stakes mental health assessments. Our findings underscore the value of multi-expert frameworks for more trustworthy AI-driven screening.

心理评估多专家模型长文本理解大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。