arXiv:2608.13787cs.AIcs.CL2026-08

让小模型学会战略性谈判,避免过度友善导致信息泄露。

From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL

论文配图:From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL
图 1 · 摘自论文原文
  • 用SocialRL直接训练模型的社会推理能力
  • 40亿参数模型在六类任务中接近甚至超越GPT-5表现
  • 基于心理理论的训练能显著提升跨领域泛化能力

AI代理越来越多地代表用户处理任务,如安排会议、比较报价和讨价还价。这些由主事者驱动的任务常使代理面对目标可能冲突的对手(其他用户的代理、卖家、招聘方)。然而,使助手显得友善的特质反而会使其成为差劲的代理人:友好的前沿模型可能未经提示就泄露用户隐私,并在遇到阻力时轻易让步。我们提出SocialRL,一种通用的训练方法,可直接训练社会推理能力,并将其应用于一个40亿参数模型,在六种场景中进行训练:出价或不出价、CaSiNo、Craigslist、求职面试、日程安排和市场交易。所有场景均采用相同方法在域内训练,且每个策略在全部六个场景中进行评估。结果表明:(1)域内训练达到前沿水平;在未见场景中,该40亿模型在每项任务上与或超过GPT-5系列表现,将基线到前沿的差距缩小73%至122%,买家初始报价低于目标的比例从3%降至78%;(2)跨域迁移遵循游戏结构:结构相似的游戏相互促进,多议题广义捐赠者几乎提升所有领域,而结构孤立的游戏则无迁移效果;(3)基于此迁移结构,采用级联强化学习和多教师在线策略蒸馏(OPD)两种策略,将各领域专长整合为单一统一的40亿模型,其在所有六种环境中的平均效用达0.627,与GPT-4.1(0.625)、GPT-5.1(0.619)和GPT-5.2(0.613)相当或更优;(4)显式心智理论框架仅在训练阶段有帮助:蒸馏心智理论轨迹而非仅动作,可提升所有环境中的效用并更好泛化,且在两类心智理论技能中,只有下一步行动预测能预测谈判结果。

原文摘要 · Abstract (English)

AI agents increasingly act on their users' behalf, handling tasks such as scheduling meetings, comparing offers, and haggling over prices. These principal-driven tasks routinely place the agent across from a counterpart (another user's agent, a seller, a recruiter) whose goals may conflict with its principal's. Yet the dispositions that make an assistant pleasant can make it a poor delegate: a friendly, helpful frontier model may disclose its principal's private information unprompted and concede at the first sign of resistance. We present SocialRL, a general recipe that trains social reasoning directly, and apply it to a 4B model across six domains: Deal-or-No-Deal, CaSiNo, Craigslist, Job Interview, Calendar, and Marketplace. Every domain is trained in-domain under the same recipe, and every policy is evaluated on all six. We find that (1) in-domain training reaches the frontier: on held-out scenarios the 4B matches or exceeds the GPT-5 family per domain, closing 73-122% of the baseline-to-frontier gap on the negotiation games, with 78% of buyer openings anchoring below target versus 3% untrained; (2) cross-domain transfer follows game structure: structurally paired games lift each other, a broad multi-issue donor lifts nearly all domains, and structurally isolated games transfer nothing; (3) guided by this transfer structure, two strategies, cascade RL and multi-teacher on-policy distillation (OPD), consolidate the per-domain specialists into a single unified 4B that reaches 0.627 average utility across all six environments, matching or exceeding GPT-4.1 (0.625), GPT-5.1 (0.619), and GPT-5.2 (0.613); (4) an explicit theory-of-mind scaffold helps only through training: distilling the ToM trace, rather than actions alone, lifts utility on every environment and generalizes better across them, and of the two ToM skills, only next-action prediction predicts negotiation outcomes.

社会推理小模型谈判强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。