arXiv:2602.16053cs.LGcs.CL2026-02

用多目标优化让AI心理咨询更贴心又安全

Multi-Objective Alignment of Language Models for Personalized Psychotherapy

  • 设计多目标对齐框架,同时优化共情、安全等六项治疗指标
  • 相比单目标方法,兼顾共情(77.6%)与安全(62.6%)表现更好
  • 适合需要个性化、高安全性AI心理助手的临床研发团队

全球超10亿人受心理健康问题困扰,但专业资源匮乏且成本高昂。现有AI疗法对齐方法独立优化目标,难以平衡患者偏好与临床安全。我们调研335名有真实心理健康经历者,收集其在治疗维度上的偏好排序,并提出基于直接偏好优化的多目标对齐框架。训练了涵盖共情、安全、主动倾听、自我驱动改变、信任/关系、患者自主性六项标准的奖励模型,系统比较多目标方法与单目标优化、监督微调及参数融合的效果。多目标DPO(MODPO)在平衡性上表现更优(共情77.6%,安全62.6%),显著优于单目标优化(共情93.6%,安全47.8%),且治疗指标比通用沟通原则高17.2%。盲评临床专家一致偏好MODPO,LLM评估者一致性接近临床医生间可靠性。

原文摘要 · Abstract (English)

Mental health disorders affect over 1 billion people worldwide, yet access to care remains limited by workforce shortages and cost constraints. While AI systems show therapeutic promise, current alignment approaches optimize objectives independently, failing to balance patient preferences with clinical safety. We survey 335 individuals with lived mental health experience to collect preference rankings across therapeutic dimensions, then develop a multi-objective alignment framework using direct preference optimization. We train reward models for six criteria -- empathy, safety, active listening, self-motivated change, trust/rapport, and patient autonomy -- and systematically compare multi-objective approaches against single-objective optimization, supervised fine-tuning, and parameter merging. Multi-objective DPO (MODPO) achieves superior balance (77.6% empathy, 62.6% safety) compared to single-objective optimization (93.6% empathy, 47.8% safety), and therapeutic criteria outperform general communication principles by 17.2%. Blinded clinician evaluation confirms MODPO is consistently preferred, with LLM-evaluator agreement comparable to inter-clinician reliability.

AI心理多目标优化大模型对齐个性化医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。