研究如何用机制防止自利智能体破坏市场,发现调解机制最抗攻击。
Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study

- 在受限社交网络中设计八种机制,对比其稳定市场能力。
- 调解机制在13.3%攻击下仍保持正收益,未导致市场崩溃。
- 适合关注AI治理、博弈机制与系统鲁棒性的研究者。
自利智能体在重复社会困境中易陷入背叛,导致贸易合作收益崩溃。本文研究何种形式机制能在无约束通信之上维持自利智能体社会的市场稳定性,并评估其对敌对攻击的韧性。通过多智能体市场模拟,18个具备互补生产能力的DeepSeek-V3大模型智能体在受限社交网络中交易以获取效用。实验分两阶段:(1) 在200轮内逐步注入恶意代理,比较八种机制,发现调解(Mediation)表现最佳;(2) 使用迭代提示优化的LLM驱动恶意代理对调解机制进行红队测试,最优攻击(v6)使诚实智能体效用下降13.3%,但无法击溃市场。调解机制在持续攻击下仍具恢复能力。定义对抗鲁棒性为机制在优化攻击下维持正向诚实效用的能力,结果显示调解机制具有鲁棒性:可被扭曲,但不可击破。
原文摘要 · Abstract (English)
Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse. This paper investigates what formal mechanisms, layered on top of unrestricted communication, are sufficient for a society of such agents to maintain market stability, and how resilient those mechanisms are to adversarial attack. We instantiate the research question as a multi-agent marketplace simulation where 18 LLM agents (DeepSeek-V3) with complementary production specialties must trade within a constrained social network to obtain utility. We conduct two experimental phases: (1) a mechanism comparison across eight conditions under progressive troll injection over 200 rounds, identifying Mediation as the top-performing mechanism; and (2) adversarial red-teaming of Mediation using iteratively prompt-optimised LLM-driven trolls, finding that the best attack (v6) reduces honest-agent utility by 13.3% but cannot collapse the market. Mediation enables recovery even under sustained adversarial pressure. We define adversarial robustness as a mechanism's ability to sustain positive honest-agent utility under optimised attack, and find that Mediation is robust: it can be bent but not broken.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。