arXiv:2602.18230cs.LGcs.AI2026-02中稿 · publication at Tra…

用可评分游戏评估大模型谈判能力,发现对比结果存疑。

[Re] Benchmarking LLM Capabilities in Negotiation through Scoreable Games

  • 基于可评分游戏构建复杂谈判评测框架。
  • 复现实验显示模型对比结果模糊,缺乏客观性。
  • 揭示信息泄露与实验设计缺陷,适合评估者参考。

大型语言模型在多智能体谈判任务中展现出巨大潜力,但该领域的评估仍面临挑战,主要源于缺乏稳健且可推广的基准测试。Abdelnabi 等人(2024)提出基于可评分游戏的谈判基准,旨在构建一个高度复杂且贴近现实的 LLM 评估框架。本文研究了该基准主张的可复现性,并深入分析其可用性与泛化能力。我们在更多模型上复现原实验,并引入额外指标以验证谈判质量与评估公平性。结果表明,尽管该基准确实具有复杂性,但模型间的比较存在模糊性,引发其客观性的质疑。此外,我们识别出实验设置中的局限,特别是信息泄露检测不充分及消融实验不彻底的问题。通过在扩展版基准上分析更广泛的模型行为,我们为潜在用户提供额外背景信息。研究强调了上下文在模型对比评估中的重要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate significant potential in multi-agent negotiation tasks, yet evaluation in this domain remains challenging due to a lack of robust and generalizable benchmarks. Abdelnabi et al. (2024) introduce a negotiation benchmark based on Scoreable Games, with the aim of developing a highly complex and realistic evaluation framework for LLMs. Our work investigates the reproducibility of claims in their benchmark, and provides a deeper understanding of its usability and generalizability. We replicate the original experiments on additional models, and introduce additional metrics to verify negotiation quality and evenness of evaluation. Our findings reveal that while the benchmark is indeed complex, model comparison is ambiguous, raising questions about its objectivity. Furthermore, we identify limitations in the experimental setup, particularly in information leakage detection and thoroughness of the ablation study. By examining and analyzing the behavior of a wider range of models on an extended version of the benchmark, we reveal insights that provide additional context to potential users. Our results highlight the importance of context in model-comparative evaluations.

大模型评估谈判模拟可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。