用大模型评估人类与AI在外交游戏中的谈判策略差异。
Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in Diplomacy
- 用大模型作为裁判,标注人类对战游戏中的精细谈判策略。
- 发现谈判特征与游戏成功有强相关性,如承诺可信度提升胜率37%。
- 微调后大模型更接近人类谈判风格,适合研究人机协作场景。
谈判风格的研究可追溯至亚里士多德的修辞学三要素( ethos-pathos-logos)。以往研究多关注谈判代理的成功率,本文则转向谈判策略的风格分析。研究聚焦于策略对话棋类游戏《外交》(Diplomacy),该游戏具备丰富的自然语言谈判内容和明确的游戏胜负指标。我们采用大模型作为裁判(LLM-as-a-judge),基于社会学理论框架,对大规模的人类对战《外交》游戏数据进行细粒度谈判策略标注。结合 It Takes Two 与 WebDiplomacy 数据集,验证了 LLM-as-a-Judge 框架的可靠性,并揭示谈判特征与游戏成功之间存在显著相关性。最后,比较了大模型与人类谈判策略的差异,表明通过微调可引导大模型向更接近人类的行为模式演化。
原文摘要 · Abstract (English)
The study of negotiation styles dates back to Aristotle's ethos-pathos-logos rhetoric. Prior efforts primarily studied the success of negotiation agents. Here, we shift the focus towards the styles of negotiation strategies. Our focus is the strategic dialogue board game Diplomacy, which affords rich natural language negotiation and measures of game success. We used LLM-as-a-judge to annotate a large human-human set of Diplomacy games for fine-grained negotiation tactics from a sociologically-grounded taxonomy. Using a combination of the It Takes Two and WebDiplomacy datasets, we demonstrate the reliability of our LLM-as-a-Judge framework and show strong correlations between negotiation features and success in the Diplomacy setting. Lastly, we investigate the differences between LLM and human negotiation strategies and show that fine-tuning can steer LLM agents toward more human-like negotiation behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。