构建对抗性金融市场评估框架,揭示现有多数AI交易代理缺乏真实市场适应能力。
TraderBench: How Robust Are AI Agents in Adversarial Capital Markets?
- 融合专家验证任务与纯绩效评分的对抗性交易模拟
- 8个模型在加密货币中表现稳定但无适应性,平均得分33
- 长思考提升知识检索,对实际交易影响微乎其微
金融领域中评估AI代理面临两大挑战:静态基准需昂贵专家标注且忽略动态决策,而基于大模型的评判引入不可控偏差。我们提出TraderBench,结合专家验证的静态任务(知识检索、分析推理)与完全基于实际绩效(夏普比率、收益、回撤)的对抗性交易模拟,彻底消除评判偏差。该框架包含两个新赛道:涵盖四种渐进式市场操纵变换的加密货币交易,以及涵盖盈亏准确率、希腊值与风险管理的期权衍生品评估。交易场景可更新市场数据以防止基准污染。在约50项任务上评估13个模型(从8B开源到前沿模型),发现:(1) 13个模型中有8个在加密货币任务中得分约33,且在对抗条件下波动小于1分,暴露了固定非自适应策略;(2) 延展思考显著提升知识检索(+26分),但对交易表现几乎无影响(加密货币仅+0.3,期权-0.1)。结果表明当前代理缺乏真正市场适应能力,凸显金融领域需基于性能的评估范式。
原文摘要 · Abstract (English)
Evaluating AI agents in finance faces two key challenges: static benchmarks require costly expert annotation yet miss the dynamic decision-making central to real-world trading, while LLM-based judges introduce uncontrolled variance on domain-specific tasks. We introduce TraderBench, a benchmark that addresses both issues. It combines expert-verified static tasks (knowledge retrieval, analytical reasoning) with adversarial trading simulations scored purely on realized performance-Sharpe ratio, returns, and drawdown-eliminating judge variance entirely. The framework features two novel tracks: crypto trading with four progressive market-manipulation transforms, and options derivatives scoring across P&L accuracy, Greeks, and risk management. Trading scenarios can be refreshed with new market data to prevent benchmark contamination. Evaluating 13 models (8B open-source to frontier) on ~50 tasks, we find: (1) 8 of 13 models score ~33 on crypto with <1-point variation across adversarial conditions, exposing fixed non-adaptive strategies; (2) extended thinking helps retrieval (+26 points) but has zero impact on trading (+0.3 crypto, -0.1 options). These findings reveal that current agents lack genuine market adaptation, underscoring the need for performance-grounded evaluation in finance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。