评估语言模型在复杂合同中的理性谈判与履约能力
Evaluating Rational Contracting in Natural Language

- 构建多轮不确定环境下的语言合同博弈框架
- 现有大模型在低不确定性下能达成有效协议,高不确定性下常失败
- 模型常为获利违反合同,缺乏合作性,适合评估AI经济行为
语言型AI代理有望拓展机器经济活动的边界。不同于仅提出报价或遵循硬编码协议,这些代理可使用自然语言进行协商与执行开放式合同。然而,现有评估多集中于单次交易或简单经济游戏,忽视了语言赋予的长期、条件性与不完整合同的丰富空间;且仅关注收益,未衡量可信合约所需品质。为此,我们提出了一个理性框架,用于指导代理在不确定的多步环境中协商与履行自然语言合同。在此框架下,我们开发了量化理性与合作行为的指标与基线。通过在ContractSim评估套件中模拟六种环境及三种供应商场景(餐饮、酒店清洁、AI托管),我们发现当前基于LLM的代理能在低环境不确定性下可靠达成协议并高效谈判。但在高不确定性下,往往无法达成可满足、高效或互利的合同。执行阶段也常表现出非合作性,即使合同易履行,仍为额外利润而违约。这些结果凸显了提升语言代理在理性与合作性合同行为上的设计空间。
原文摘要 · Abstract (English)
The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols, such agents can be used to negotiate and execute agreements in open-ended natural language. However, most evaluations of these abilities have focused on one-off exchanges or simple economic games, leaving open the rich space of time-extended, contingent, and incomplete contracts made expressible by language; they also focus on raw profit, without measuring the qualities required for trustworthy contracting. We address this by formulating a rational framework for how agents should negotiate and perform natural language contracts in uncertain multi-step environments. Within this framework, we develop metrics and baselines for quantifying rational and cooperative play. To evaluate how agents perform at such contracting, we instantiate our framework in ContractSim, an evaluation suite where two players negotiate and execute a multi-turn supplier contract under environmental and inter-player uncertainty. Across six environments and three supplier settings (catering, hotel cleaning, and AI hosting) we find that current LLM-based agents reach agreement reliably, and negotiate efficient contracts when environmental uncertainty is low. However, under high uncertainty, they often fail to negotiate satisfiable, efficient, or mutually beneficial contracts. They are also frequently uncooperative when executing contracts, violating contract terms for additional profit even when contracts are easy to satisfy. These findings highlight room for improvement in the design of language agents that can negotiate, interpret, and execute contracts both rationally and cooperatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。