用动态策略注入提升语言代理社交能力,效果超GPT-4
SOTOPIA-$Ω$: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social Agents
- 通过谈判理论注入多步推理策略,动态构建高质量对话训练数据
- 7B模型在社交目标达成上超越GPT-4,S-IF指标显著提升
- 适合研究社交智能、人机交互的团队,尤其关注策略学习与评估
尽管人类拥有丰富的社交策略,但将其迁移到社交代理的研究仍十分有限。本文提出的 SOTOPIA-Ω 框架旨在弥补这一空白,重点提升语言代理的社交能力。该框架将谈判理论启发的多步推理策略及两种简单直接策略动态注入专家代理,从而自动化构建高质量社交对话训练语料库。同时,提出社交指令遵循(S-IF)概念,并设计两种新评估指标以补充社交能力评价。实验表明,多个7B规模模型在高质量语料上训练后,不仅显著优于专家代理(GPT-4)在实现社交目标上的表现,且在S-IF任务中性能大幅提升。分析与变体实验验证了动态构建的优势,尤其能有效打破代理长期僵局。
原文摘要 · Abstract (English)
Despite the abundance of prior social strategies possessed by humans, there remains a paucity of research dedicated to their transfer and integration into social agents. Our proposed SOTOPIA-$Ω$ framework aims to address and bridge this gap, with a particular focus on enhancing the social capabilities of language agents. This framework dynamically injects multi-step reasoning strategies inspired by negotiation theory and two simple direct strategies into expert agents, thereby automating the construction of a high-quality social dialogue training corpus. Additionally, we introduce the concept of Social Instruction Following (S-IF) and propose two new S-IF evaluation metrics that complement social capability. We demonstrate that several 7B models trained on high-quality corpus not only significantly surpass the expert agent (GPT-4) in achieving social goals but also enhance S-IF performance. Analysis and variant experiments validate the advantages of dynamic construction, which can especially break the agent's prolonged deadlock.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。