arXiv:2609.05663cs.AIcs.CE2026-09

真实交易中大模型代理如何表现?六月大规模观测揭示行为真相。

What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets

论文配图:What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
图 1 · 摘自论文原文
  • 通过两套生产级系统持续监测,记录超750万次调用与30万次链上操作。
  • 杠杆行为由系统设置决定而非策略文本,60%波动来自个体代理差异。
  • 多数高收益位置最终亏损,仅少数能稳定获利,适合研究智能合约风险者看。

本文呈现了两个同源系统中自主语言模型交易代理在生产环境下的连续、大规模观测记录:DX Terminal Pro(3,505个用户资助的钱包,在Base memecoin市场交易ETH,持续21天,2026年2月至3月)与DXAP实时测试舰队(500至599个用户创建的代理,全历史数据,91至117个并发活跃,交易Hyperliquid永续合约,2026年6月至8月)。记录跨度约六个月,涵盖约750万次单模型调用及约30万次链上动作,另有231,638次多工具协作产生14,596次成交。四个核心发现:第一,运行层设计远比策略文本影响行为——风险滑块每提升一级,杠杆增加0.425倍;代理固定效应解释60%方差;排行榜边界存在因果选择效应(前3名截断点处回报率提升1.75倍)。第二,仓位大小无视波动率:所有波动率六分位区间中位杠杆均为5.0倍;一个姿态滑块单元(占总持仓11%)承担62%的清算。第三,代理几乎未捕获其达到的利润:43.2%的头寸在24小时内出现至少+300基点有利价差,但49.3%在此后以负回报平仓;机械括号策略可恢复+39.0基点/头寸。第四,两舰队均无方向性优势:DXAP舰队不盈利,回测胜率41%低于匹配的Hyperliquid零售基准(50%)。对416个生产场景进行配对重播测试显示,前沿模型间决策质量在当前时间窗口内无统计差异,但模型族间选择稳定性差异显著。所有结论经日聚类推断、置换零假设与统一手续费重述验证,论文末尾附17条方法论准则,并包含自我撤稿记录。

原文摘要 · Abstract (English)

We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.

量化交易LLM应用生产实测区块链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。