双智能体协同训练提升AI健康教练效果
Dual-Agent Co-Training for Health Coaching via Implicit Adversarial Preference Optimization

- 双向交互式训练教练与用户模拟器,打破单向优化局限
- 用多维大模型判官筛选优质对话对,提升教练响应质量
- 隐式对抗训练让教练更懂用户心理,适合医疗对话研究者
基于动机访谈的健康辅导是改善心理健康和促进行为改变的有效方法。然而,受过训练的人类教练稀缺且辅导成本高昂,导致许多需要帮助的人难以获得支持。这推动了人工智能健康教练的发展,以实现可扩展、低成本的辅助。现有方法通常只优化互动的一方:要么训练对话代理针对固定客户环境,要么训练客户模拟器针对固定助手。这种单边设置限制了互动空间的探索,可能在发展目标代理能力及突破性能边界方面效率低下。本文提出一种双智能体框架,通过交互式协同训练健康教练代理与客户模拟器。教练采用基于偏好优化(DPO)的方法,利用多维度大模型判官识别帕累托占优的回应对进行优化;同时,客户模拟器通过反转这些偏好进行对抗性训练,形成隐式对抗训练动态。我们进一步证明该协同训练过程具有自然的随机博弈解释。大量实验表明,该方法在多个关键维度上显著提升了辅导质量。
原文摘要 · Abstract (English)
Motivational-interviewing-based health coaching is an effective approach for improving mental health and promoting healthy behavior change. However, the scarcity of trained human coaches and the high cost of coaching services make such support inaccessible to many people who could benefit from it. This motivates the development of AI health coaches that can provide scalable and affordable support. Existing methods typically optimize only one side of the interaction: they either train a dialogue agent against a fixed client environment or train a client simulator against a fixed assistant. This one-sided setup can limit exploration of the interaction space and may be inefficient at developing the capabilities required by the target agent and pushing its performance boundaries. In this paper, we propose a dual-agent framework that interactively co-trains both the health coach agent and the client simulator. The coach is optimized with DPO using Pareto-dominant response pairs identified by a multi-dimensional LLM judge. In turn, the client is trained adversarially by reversing these preferences, inducing an implicit adversarial training dynamic. We further show that this co-training process admits a natural stochastic-game interpretation. Extensive experiments demonstrate that our method effectively improves coaching quality across several important dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。