arXiv:2607.28814cs.CLcs.AI2026-07

用偏好优化让大模型更会倾听,但可能牺牲改变目标的决心。

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing

论文配图:Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing
图 1 · 摘自论文原文
  • 通过惩罚对抗或妥协行为,训练模型更顺应客户情绪
  • 惩罚对抗可显著降低目标坚持度,但情感契合度提升因模型而异
  • 纯提示控制能提升契合度而不损目标,说明代价来自优化过程

在动机访谈中,客户维持现状的言辞需要咨询师‘顺应阻力’,否则可能陷入妥协(放弃改变目标以保关系)或对抗(强行说服)。本文基于MITI评分体系,构建双轴评估框架:目标坚持(GP)与关系契合(RA),形成四象限。利用专家标注的AnnoMI语料库,建立话题不重叠的直接偏好优化数据集,仅在拒绝哪种失败上不同,并采用同策略负样本。通过自动评判器(经专家标签验证并由训练人员复核),在防火墙内对三组对齐指令模型(涵盖Qwen与Llama系列)进行盲测,比较成对胜率。结果表明:惩罚对抗行为能稳定降低目标坚持度至基线以下,且在所有基线与种子运行中均成立;契合度增益则依赖于基线模型,在其中两个基线中存在,第三个则无。惩罚妥协无效,因这些模型本就极少出现该行为,权衡取决于各基线的失败特征。仅使用提示的对照实验提升了契合度而未损失目标坚持,说明成本源于优化过程本身而非契合度提升。

原文摘要 · Abstract (English)

In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to preserve rapport) or confrontation (arguing or directing, overriding the client's autonomy). We introduce a two-axis evaluation of counselor responses, anchored in the Motivational Interviewing Treatment Integrity (MITI) code, Goal Persistence (GP) and Relational Attunement (RA), yielding a four-quadrant framing in which rolling with resistance is high on both, and we ask whether penalizing one failure through preference optimization teaches rolling with resistance or provokes its opposite. From the expert-annotated AnnoMI corpus we build topic-disjoint Direct Preference Optimization data whose preference sets differ only in which failure is rejected, using on-policy negatives. An automatic judge, validated against AnnoMI's expert labels and rechecked by trained human coders, scores blind pairwise win-rates against each base under a firewall in which disjoint model families generate, label, and judge. Across three aligned instruction models spanning the Qwen and Llama families, penalizing confrontation reliably lowers goal persistence below parity, on every base and in every seed run, a robust cost, whereas the attunement gain is base-dependent, present on two of the three bases but absent on the third. Penalizing capitulation is inert, because these models rarely capitulate on-policy, so the trade is gated by each base's failure profile. A prompt-only control raises attunement without the goal-persistence cost, locating the cost in the optimization rather than in attunement itself.

动机访谈大模型训练偏好优化人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。