用实时市场环境测试大模型交易决策能力,发现排行榜高分不等于赚钱能力强。
LiveTradeBench: Seeking Real-World Alpha with Large Language Models
- 构建实时流式市场数据环境,模拟真实不确定性下的投资决策。
- 21个大模型在50天内表现各异,部分模型能有效利用实时信息调整策略。
- 适用于评估大模型在动态环境中的持续适应能力和风险控制水平。
大语言模型在各类静态基准测试中表现优异,但这些测试缺乏真实市场的动态性与不确定性,难以评估其在不确定条件下的决策能力。为此,我们提出LiveTradeBench,一个面向真实演进市场的实时交易评估环境。该平台遵循三大设计原则:(i) 实时流式传输市场价格与新闻数据,避免离线回测依赖和信息泄露,捕捉真实不确定性;(ii) 采用组合管理抽象,将控制从单一资产扩展到多资产配置,融合风险管理和跨资产推理;(iii) 在结构不同的市场(美国股票与Polymarket预测市场)上进行多市场评估,涵盖波动性、流动性与信息流动差异。每个时间步,代理观察价格、新闻与持仓,输出风险与收益平衡的资产分配比例。通过50天的实时评估,对21个大模型家族进行测试,结果表明:(1) 高LMArena评分不代表更优交易表现;(2) 模型表现出反映风险偏好与推理动态的差异化组合风格;(3) 部分模型能有效利用实时信号调整决策。研究揭示了静态评估与真实世界能力之间的差距,呼吁开发能测试序列决策与持续适应性的新基准。
原文摘要 · Abstract (English)
Large language models (LLMs) achieve strong performance across benchmarks--from knowledge quizzes and math reasoning to web-agent tasks--but these tests occur in static settings, lacking real dynamics and uncertainty. Consequently, they evaluate isolated reasoning or problem-solving rather than decision-making under uncertainty. To address this, we introduce LiveTradeBench, a live trading environment for evaluating LLM agents in realistic and evolving markets. LiveTradeBench follows three design principles: (i) Live data streaming of market prices and news, eliminating dependence on offline backtesting and preventing information leakage while capturing real-time uncertainty; (ii) a portfolio-management abstraction that extends control from single-asset actions to multi-asset allocation, integrating risk management and cross-asset reasoning; and (iii) multi-market evaluation across structurally distinct environments--U.S. stocks and Polymarket prediction markets--differing in volatility, liquidity, and information flow. At each step, an agent observes prices, news, and its portfolio, then outputs percentage allocations that balance risk and return. Using LiveTradeBench, we run 50-day live evaluations of 21 LLMs across families. Results show that (1) high LMArena scores do not imply superior trading outcomes; (2) models display distinct portfolio styles reflecting risk appetite and reasoning dynamics; and (3) some LLMs effectively leverage live signals to adapt decisions. These findings expose a gap between static evaluation and real-world competence, motivating benchmarks that test sequential decision making and consistency under live uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。