arXiv:2601.13481cs.AI2026-01

用多智能体优化提示词,提升心理情绪诊断的准确性和稳定性。

Towards Efficient and Robust Linguistic Emotion Diagnosis for Mental Health via Multi-Agent Instruction Refinement

  • 构建多智能体系统,自动迭代优化诊断提示词
  • 在多个心理评估数据集上提升诊断准确率与鲁棒性
  • 适合临床心理、AI医疗等需要高可靠性情绪分析的场景

抑郁、焦虑及创伤相关情绪的语言表达广泛存在于病历记录、心理咨询对话和在线心理健康社区中,其精准识别对临床分诊、风险评估和及时干预至关重要。尽管大语言模型在情绪分析任务中展现出强大泛化能力,但在高风险、强上下文依赖的医疗场景中,其诊断可靠性仍高度依赖提示词设计。现有方法面临两大挑战:多种情绪共病导致预测复杂,以及对临床关键线索的探索效率低下。为此,我们提出APOLO(自动化提示优化用于语言情绪诊断)框架,系统探索更广、更细粒度的提示空间以提升诊断效率与鲁棒性。APOLO将指令优化建模为部分可观测马尔可夫决策过程,采用包含规划者、教师、批评者、学生和目标角色的多智能体协作机制。在闭环框架中,规划者设定优化路径,教师-批评者-学生代理迭代优化提示以增强推理稳定性与有效性,目标代理根据性能评估决定是否继续优化。实验表明,APOLO在领域特定和分层基准上均持续提升诊断准确率与鲁棒性,展示出可信大模型在心理健康领域的可扩展、通用范式。

原文摘要 · Abstract (English)

Linguistic expressions of emotions such as depression, anxiety, and trauma-related states are pervasive in clinical notes, counseling dialogues, and online mental health communities, and accurate recognition of these emotions is essential for clinical triage, risk assessment, and timely intervention. Although large language models (LLMs) have demonstrated strong generalization ability in emotion analysis tasks, their diagnostic reliability in high-stakes, context-intensive medical settings remains highly sensitive to prompt design. Moreover, existing methods face two key challenges: emotional comorbidity, in which multiple intertwined emotional states complicate prediction, and inefficient exploration of clinically relevant cues. To address these challenges, we propose APOLO (Automated Prompt Optimization for Linguistic Emotion Diagnosis), a framework that systematically explores a broader and finer-grained prompt space to improve diagnostic efficiency and robustness. APOLO formulates instruction refinement as a Partially Observable Markov Decision Process and adopts a multi-agent collaboration mechanism involving Planner, Teacher, Critic, Student, and Target roles. Within this closed-loop framework, the Planner defines an optimization trajectory, while the Teacher-Critic-Student agents iteratively refine prompts to enhance reasoning stability and effectiveness, and the Target agent determines whether to continue optimization based on performance evaluation. Experimental results show that APOLO consistently improves diagnostic accuracy and robustness across domain-specific and stratified benchmarks, demonstrating a scalable and generalizable paradigm for trustworthy LLM applications in mental healthcare.

情绪诊断多智能体LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。