arXiv:2606.13176cs.AI2026-06被引 1

用认知过程模拟增强大模型心理评估推理能力

Mental-R1: Aligning LLM Reasoning for Mental Health Assessment

论文配图:Mental-R1: Aligning LLM Reasoning for Mental Health Assessment
图 1 · 摘自论文原文
  • 设计分阶段熵正则机制,模仿人类从不确定到确定的思考过程
  • 在8个数据集上提升平均10.4个百分点的加权F1分数
  • 适合需要可解释推理的心理健康评估研究者使用

焦虑、抑郁和自杀等心理健康问题仍是全球重大挑战,及时准确的评估对有效干预至关重要。尽管大语言模型已被用于心理评估,但现有通用后训练方法未能对齐人类评估的认知过程,可能导致不可靠的推理结果。为此,我们提出认知相对策略优化(CRPO),一种专为心理健康领域设计的强化学习框架。CRPO通过引入阶段依赖的不确定性建模,在策略优化中扩展了组相对策略优化。具体而言,我们设计分阶段熵正则化机制,鼓励早期推理阶段广泛探索,后期逐步强化自信决策,模拟人类从不确定到确定的认知转变。此外,受认知评价理论启发,我们形式化了认知推理阶段,实现理论引导的可解释推断。在8个心理健康数据集上的实验表明,CRPO相较于最佳强化学习基线,平均提升10.4个百分点的加权F1分数。同时,经CRPO训练的Mental-R1模型在推理密集型案例中显著优于现有大语言模型,表明CRPO有效提升了心理评估中的推理能力。

原文摘要 · Abstract (English)

Mental health problems such as anxiety, depression, and suicide remain urgent global challenges, where timely and accurate assessment is critical for effective intervention. Recently, large language models have been explored for mental health assessment. However, existing general-purpose post-training methods do not align with the cognitive processes of human assessment, which may lead to unreliable reasoning outcomes. To bridge this gap, we propose Cognitive Relative Policy Optimization (CRPO), a reinforcement learning framework tailored for the mental health domain. CRPO extends group relative policy optimization by integrating stage-dependent uncertainty modeling into the policy optimization process. Specifically, we introduce a stage-wise entropy regularization mechanism that encourages broad exploration in early reasoning phases and progressively enforces confident decision-making in later stages, mimicking the human cognitive shift from uncertainty to certainty. In addition, inspired by cognitive appraisal theory, we formalize cognitive reasoning stages, thereby guiding theory-grounded interpretable inference. Experiments on 8 mental health datasets show that CRPO achieves an average improvement of 10.4 percentage points in weighted F1-score over the best reinforcement learning baseline. Furthermore, the CRPO-trained model Mental-R1 demonstrates clear advantages compared with existing large language models on reasoning-intensive cases, suggesting that CRPO enhances reasoning capabilities for mental health assessment.

心理评估推理增强强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。