arXiv:2605.26785cs.CLcs.AI2026-05被引 1

让大模型学会战略性使用情绪,提升谈判表现

EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation

论文配图:EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
图 1 · 摘自论文原文
  • 分离情绪选择与表达,用强化学习选情绪、LoRA微调表达方式
  • 在4个高风险谈判场景中,新方法达成最高收益,超越基线模型
  • 可离线训练,适配新对手和场景,避免真实谈判成本

后训练的大语言模型通常优化为符合人类偏好,表现为安全、礼貌和对话得体。但在对抗性谈判中,这种对齐可能成为弱点:情感化语言可能引导代理偏向对方利益。基于GoEmotions的情感提示实验表明,情绪显著改变谈判结果,说明情绪是策略性行动渠道而非表面风格。为此,我们提出EmoDistill——一种从离线代理间交互中蒸馏情绪谈判技能的框架。EmoDistill将情绪策略分解为情绪选择与表达:隐式Q学习(IQL)选择‘表达何种情绪’,基于LoRA的策略通过监督微调(SFT)与裁判策略优化(JPO)学习‘如何表达’。在四个情绪敏感、高风险谈判领域中,经EmoDistill训练的SLM策略取得最高效用,优于基础SLM/LLM及仅用IQL的情绪选择模型。消融实验表明情绪条件至关重要,迁移实验显示跨领域、未见对手及训练对战中的泛化能力。总体而言,EmoDistill从离线交互中学习技能,无需训练时进行实际谈判。

原文摘要 · Abstract (English)

Post-trained LLMs are often optimized to align responses with human preferences, making them safe, polite, and conversationally appropriate. In adversarial negotiation, however, this alignment can become a vulnerability: emotionally framed language may steer agents toward the counterparty's interests. Using GoEmotions-based affective prompting, we show that emotion substantially shifts negotiation outcomes, suggesting that emotion is a strategic action channel rather than a surface style. Thus, we introduce \textbf{EmoDistill}, an offline framework for distilling emotional negotiation skills into language model agents. EmoDistill decomposes emotional strategy into emotion selection and emotion expression: an Implicit Q-Learning (IQL) selector learns \emph{which} emotion to express, while a Low-Rank Adaptation (LoRA)-based policy learns \emph{how} to express it through Supervised Fine-Tuning (SFT) and Judge Policy Optimization (JPO). Across four emotion-sensitive, high-stakes negotiation domains, SLM policies trained under the EmoDistill framework achieve the highest utility, outperforming vanilla SLM/LLM baselines and IQL-only emotion selection. Ablations show that emotion conditioning is essential, and transfer studies demonstrate generalization across domains, unseen counterparties, and trained-vs-trained tournaments. Overall, EmoDistill learns skills from offline agent-to-agent interactions, avoiding costly online negotiation during training.

谈判智能情绪建模离线训练强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。