验证大模型谈判中合作与竞争机制,发现小模型表现差但大模型可媲美闭源模型。
Reproducibility Study of Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
- 用1.5B到70B参数的开源模型复现并扩展原研究,测试多模型协作谈判效果。
- 小模型(<10B)难以保持格式一致和连贯输出,大模型接近闭源性能。
- 单智能体方案常可替代多智能体谈判,挑战了交互必要性的假设。
本文对《合作、竞争与恶意:大语言模型-利益相关者互动谈判》进行了可复现性研究与拓展。我们使用一系列开源模型(1.5B–70B参数)和GPT-4o Mini验证原始结论,并提出若干新贡献:分析游戏的帕累托前沿,设计无需通信的基线以检验无交互谈判可行性,评估近期小型语言模型的表现,分析模型响应中的结构信息泄露,以及实现不平等度量以评估谈判公平性。结果表明,小于10B参数的小模型在格式遵循和连贯性上表现不佳,而更大规模的开源模型可逼近专有模型性能。此外,在许多场景下,单智能体方法可达到与多智能体谈判相当的效果,质疑了任务成功必须依赖智能体间交互的假设。本研究还揭示了基于大模型谈判系统的可访问性、公平性、环境影响与隐私问题。
原文摘要 · Abstract (English)
This paper presents a reproducibility study and extension of "Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation." We validate the original findings using a range of open-weight models (1.5B-70B parameters) and GPT-4o Mini while introducing several novel contributions. We analyze the Pareto front of the games, propose a communication-free baseline to test whether successful negotiations are possible without agent interaction, evaluate recent small language models' performance, analyze structural information leakage in model responses, and implement an inequality metric to assess negotiation fairness. Our results demonstrate that smaller models (<10B parameters) struggle with format adherence and coherent responses, but larger open-weight models can approach proprietary model performance. Additionally, in many scenarios, single-agent approaches can achieve comparable results to multi-agent negotiations, challenging assumptions about the necessity of agent communication to perform well on the benchmark. This work also provides insights into the accessibility, fairness, environmental impact, and privacy considerations of LLM-based negotiation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。