arXiv:2512.02261cs.AI2025-12被引 3

测试大模型交易代理在扰动下的可靠性,发现微小干扰可引发重大损失。

TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?

  • 构建统一框架,在真实美股数据上闭环回测
  • 单点扰动导致极端持仓集中与巨额回撤
  • 适合关注AI金融风险的研究者与从业者

基于大语言模型的交易代理正被广泛应用于实际金融市场,执行自主分析与交易。然而,其在对抗性或故障条件下的可靠性与鲁棒性尚未充分研究,而这些系统运行于高风险、不可逆的金融环境。本文提出TradeTrap,一个统一的评估框架,用于系统性地压力测试自适应与程序化自主交易代理。该框架针对交易代理的四大核心组件——市场感知、策略制定、组合与账本管理、交易执行——在受控系统级扰动下进行鲁棒性评估。所有评估均在相同初始条件下,基于真实美国股市历史数据的闭环回测环境中进行,确保跨代理与攻击方式的公平、可复现比较。大量实验表明,单一组件的微小扰动可在代理决策循环中传播,导致极端持仓集中、失控敞口及大幅组合回撤,证明当前自主交易代理在系统层面可能被系统性误导。代码已开源:https://github.com/Yanlewen/TradeTrap。

原文摘要 · Abstract (English)

LLM-based trading agents are increasingly deployed in real-world financial markets to perform autonomous analysis and execution. However, their reliability and robustness under adversarial or faulty conditions remain largely unexamined, despite operating in high-risk, irreversible financial environments. We propose TradeTrap, a unified evaluation framework for systematically stress-testing both adaptive and procedural autonomous trading agents. TradeTrap targets four core components of autonomous trading agents: market intelligence, strategy formulation, portfolio and ledger handling, and trade execution, and evaluates their robustness under controlled system-level perturbations. All evaluations are conducted in a closed-loop historical backtesting setting on real US equity market data with identical initial conditions, enabling fair and reproducible comparisons across agents and attacks. Extensive experiments show that small perturbations at a single component can propagate through the agent decision loop and induce extreme concentration, runaway exposure, and large portfolio drawdowns across both agent types, demonstrating that current autonomous trading agents can be systematically misled at the system level. Our code is available at https://github.com/Yanlewen/TradeTrap.

交易代理鲁棒性测试LLM应用金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。