越强的模型反而让金融系统更危险,因它们行为越来越相似。
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets

- 用模拟交易员测试不同能力模型的行为相关性
- 模型越强,动作越趋同,市场风险不降反升
- 适合关注AI系统风险与金融稳定的研究者
大语言模型正被大规模部署于金融、内容审核和招聘等关键系统。我们发现提升单个模型能力可能反而恶化系统级结果。假设共享训练与架构导致更强大的LLM行为趋于一致,产生不可分散的风险。我们提出一个通用框架,证明这种相关性会形成非可分散风险下限,并通过包含不同通用能力水平的LLM交易员的代理模型仿真,在金融市场中验证其预测:(1) 前沿模型表现出显著相关的行动,且相关性随能力提升而增加;(2) 当共享推理准确时,增加代理参与可降低市场风险;(3) 当代理共享错误信息环境时,相同的相关行为成为负担。这些结果揭示了能力悖论:提升个体模型并不必然带来更好的系统表现。这一机制在其他领域是否成立,尚待实证检验。
原文摘要 · Abstract (English)
Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes. We hypothesize that shared training and architectures can lead more capable LLMs to behave more similarly, creating correlated actions that do not diversify away. We develop a general framework showing how this correlation creates a non-diversifiable risk floor and test its predictions in financial markets using an agent-based simulation with LLM traders of varying general-purpose capability. We find that: (1) frontier LLMs exhibit significantly correlated behavior that increases with capability; (2) when their shared reasoning is accurate, increasing agent participation reduces market-level risk; and (3) when agents share a common misinformation environment, the same correlated behavior becomes a liability. Together, these results identify a capability paradox: improving individual models does not necessarily produce better system-level outcomes. Whether the same dynamics arise in other domains is an open empirical question.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。