真实资本下,语言模型代理的可靠性依赖操作层设计而非模型本身。
Operating-Layer Controls for Onchain Language-Model Agents Under Real Capital

- 构建操作层控制体系,整合提示编译、策略验证与执行守卫。
- 99.9%交易结算成功率,78%资本部署率,减少虚假规则与费用僵局。
- 适合关注智能合约代理可靠性的开发者与量化研究者。
我们研究了在真实资本环境下,将用户指令转化为经验证工具动作的自主语言模型代理的可靠性。实验基于DX Terminal Pro系统,在21天内部署了3,505个由用户资助的代理,在受限的链上市场中交易真实ETH。用户通过结构化控制和自然语言策略配置资金池,但仅代理可执行常规买卖操作。系统共产生750万次代理调用,约30万次链上动作,交易额达约2000万美元,部署超5000 ETH,生成约700亿推理令牌,政策有效提交交易的结算成功率达99.9%。长期运行的代理完成数千次连续决策,包括6000+次提示-状态-动作循环,形成从用户指令到提示生成、推理、验证、投资组合状态及结算的完整追踪数据。可靠性并非来自基础模型,而是源于围绕模型的操作层:提示编译、类型化控制、策略验证、执行守卫、记忆设计与追踪可观测性。预发布测试揭示了文本基准常忽略的故障,如虚构交易规则、费用僵局、数值锚定、节奏交易和误读代币经济。针对性优化使虚构卖出规则占比从57%降至3%,费用驱动观察从32.5%降至10%以下,受影响测试组的资本部署率从42.9%提升至78.0%。研究证明,资本管理代理应贯穿用户指令到提示、验证动作与结算的全链路评估。
原文摘要 · Abstract (English)
We study reliability in autonomous language-model agents that translate user mandates into validated tool actions under real capital. The setting is DX Terminal Pro, a 21-day deployment in which 3,505 user-funded agents traded real ETH in a bounded onchain market. Users configured vaults through structured controls and natural-language strategies, but only agents could choose normal buy/sell trades. The system produced 7.5M agent invocations, roughly 300K onchain actions, about $20M in volume, more than 5,000 ETH deployed, roughly 70B inference tokens, and 99.9% settlement success for policy-valid submitted transactions. Long-running agents accumulated thousands of sequential decisions, including 6,000+ prompt-state-action cycles for continuously active agents, yielding a large-scale trace from user mandate to rendered prompt, reasoning, validation, portfolio state, and settlement. Reliability did not come from the base model alone; it emerged from the operating layer around the model: prompt compilation, typed controls, policy validation, execution guards, memory design, and trace-level observability. Pre-launch testing exposed failures that text-only benchmarks rarely measure, including fabricated trading rules, fee paralysis, numeric anchoring, cadence trading, and misread tokenomics. Targeted harness changes reduced fabricated sell rules from 57% to 3%, reduced fee-led observations from 32.5% to below 10%, and increased capital deployment from 42.9% to 78.0% in an affected test population. We show that capital-managing agents should be evaluated across the full path from user mandate to prompt, validated action, and settlement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。