用人类偏好反馈提升大模型谈判能力,让机器更懂人心。
MERIT Feedback Elicits Better Bargaining in LLM Negotiators
- 基于效用理论设计反馈机制,引导模型学习人类偏好
- 在9种复杂场景中,策略契合度显著提升,对手感知更强
- 适合研究人机博弈、智能谈判系统的设计者
谈判常被视为逻辑过程而非艺术或直觉,但大语言模型(LLMs)仍因战略深度不足且难以适应复杂人类因素而表现不佳。现有基准未能充分反映这一局限。为此,我们提出一种以效用反馈为核心的框架:(i) 构建新基准 AgoraBench,涵盖九种挑战性场景(如欺骗、垄断),支持多样化策略建模;(ii) 设计符合人类偏好的经济化度量指标,通过代理效用、议价能力与获取比率等隐式衡量谈判与人类偏好的一致性;(iii) 构建基于人类偏好的数据集及训练流程,结合提示与微调增强模型谈判能力。实证结果表明,基线模型策略常偏离人类偏好,而我们的方法显著提升表现,实现更深的战略行为与更强的对手感知。
原文摘要 · Abstract (English)
Bargaining is often regarded as a logical arena rather than an art or a matter of intuition, yet Large Language Models (LLMs) still struggle to navigate it due to limited strategic depth and difficulty adapting to complex human factors. Current benchmarks rarely capture this limitation. To bridge this gap, we present a utility feedback centric framework. Our contributions are: (i) AgoraBench, a new benchmark spanning nine challenging settings (e.g., deception, monopoly) that supports diverse strategy modeling; (ii) human-aligned, economically grounded metrics derived from utility theory. This is operationalized via agent utility, negotiation power, and acquisition ratio that implicitly measure how well the negotiation aligns with human preference and (iii) a human preference grounded dataset with learning pipeline that strengthens LLMs' bargaining ability through both prompting and finetuning. Empirical results indicate that baseline LLM strategies often diverge from human preferences, while our mechanism substantially improves negotiation performance, yielding deeper strategic behavior and stronger opponent awareness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。