用效用反馈提升大模型讨价还价能力,更贴近真实人类互动。
LLM Agents for Bargaining with Utility-based Feedback
- 基于效用理论设计反馈机制,引导模型理解对手意图。
- 在六种复杂场景中,模型策略与人类偏好差距显著缩小。
- 适合研究人机博弈、策略推理或对齐评估的学者参考。
讨价还价是现实交互中的关键环节,但大语言模型(LLMs)因战略深度不足和对复杂人类因素适应性差而面临挑战。现有基准难以反映真实世界的复杂性。为此,我们提出一个以效用为基础的反馈框架,推动LLMs在真实讨价还价中的能力提升。贡献包括:(1)BargainArena,一个包含六种复杂场景(如欺骗行为、垄断)的新基准数据集,支持多样策略建模;(2)受效用理论启发的人类对齐评估指标,结合代理效用与谈判力,隐式促进对手感知推理(OAR);(3)结构化反馈机制,使LLMs可迭代优化策略,并可与上下文学习(ICL)提示协同,尤其能增强显式设计的OAR提示效果。实验表明,模型原有策略常偏离人类偏好,而我们的反馈机制显著提升表现,带来更深的战略思考与对手意识。
原文摘要 · Abstract (English)
Bargaining, a critical aspect of real-world interactions, presents challenges for large language models (LLMs) due to limitations in strategic depth and adaptation to complex human factors. Existing benchmarks often fail to capture this real-world complexity. To address this and enhance LLM capabilities in realistic bargaining, we introduce a comprehensive framework centered on utility-based feedback. Our contributions are threefold: (1) BargainArena, a novel benchmark dataset with six intricate scenarios (e.g., deceptive practices, monopolies) to facilitate diverse strategy modeling; (2) human-aligned, economically-grounded evaluation metrics inspired by utility theory, incorporating agent utility and negotiation power, which implicitly reflect and promote opponent-aware reasoning (OAR); and (3) a structured feedback mechanism enabling LLMs to iteratively refine their bargaining strategies. This mechanism can positively collaborate with in-context learning (ICL) prompts, including those explicitly designed to foster OAR. Experimental results show that LLMs often exhibit negotiation strategies misaligned with human preferences, and that our structured feedback mechanism significantly improves their performance, yielding deeper strategic and opponent-aware reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。