用真实金融市场亏损迫使大模型代理放弃幻觉,实现可靠对齐。
OOM-RL: Out-of-Money Reinforcement Learning Market-Driven Alignment for LLM-Based Multi-Agent Systems

- 将代理投入实时金融市场,以资本耗尽为不可篡改的负向激励。
- 最终系统实现年化夏普比率2.06,代码覆盖率稳定在95%以上。
- 适合高风险场景下需客观约束的自主智能体系统设计。
自主软件工程中的多智能体系统(MAS)对齐受限于评估者认知不确定性。当前范式如基于人类反馈的强化学习(RLHF)和基于AI反馈的强化学习(RLAIF)常导致模型谄媚,而基于执行环境的方法则易受无约束代理的对抗性“测试规避”影响。本文提出一种目标对齐新范式:外生资金强化学习(OOM-RL)。通过将代理部署于非平稳、高摩擦的真实金融市,利用关键资本耗尽作为不可破解的负梯度。为期20个月的纵向实证研究(2024年7月—2026年2月)记录了系统从高周转、谄媚型基线演进为稳健、流动性敏感的架构。结果表明,财务损失的不可回避本体后果迫使MAS摒弃过拟合幻觉,转而采用严格测试驱动的智能体工作流(STDAW),其依赖确定性验证的≥95%代码覆盖率约束矩阵,实现拜占庭式单向状态锁定(RO-Lock)。早期迭代虽有严重执行衰减,但最终的OOM-RL对齐系统在成熟阶段达成了稳定的均衡,年化夏普比率达2.06。结论表明,以严格的经济惩罚替代主观人类偏好,可为高风险真实环境中的自主代理提供稳健对齐方法,为计算计费作为客观物理约束的通用范式奠定基础。
原文摘要 · Abstract (English)
The alignment of Multi-Agent Systems (MAS) for autonomous software engineering is constrained by evaluator epistemic uncertainty. Current paradigms, such as Reinforcement Learning from Human Feedback (RLHF) and AI Feedback (RLAIF), frequently induce model sycophancy, while execution-based environments suffer from adversarial "Test Evasion" by unconstrained agents. In this paper, we introduce an objective alignment paradigm: \textbf{Out-of-Money Reinforcement Learning (OOM-RL)}. By deploying agents into the non-stationary, high-friction reality of live financial markets, we utilize critical capital depletion as an un-hackable negative gradient. Our longitudinal 20-month empirical study (July 2024 -- February 2026) chronicles the system's evolution from a high-turnover, sycophantic baseline to a robust, liquidity-aware architecture. We demonstrate that the undeniable ontological consequences of financial loss forced the MAS to abandon overfitted hallucinations in favor of the \textbf{Strict Test-Driven Agentic Workflow (STDAW)}, which enforces a Byzantine-inspired uni-directional state lock (RO-Lock) anchored to a deterministically verified $\geq 95\%$ code coverage constraint matrix. Our results show that while early iterations suffered severe execution decay, the final OOM-RL-aligned system achieved a stable equilibrium with an annualized Sharpe ratio of 2.06 in its mature phase. We conclude that substituting subjective human preference with rigorous economic penalties provides a robust methodology for aligning autonomous agents in high-stakes, real-world environments, laying the groundwork for generalized paradigms where computational billing acts as an objective physical constraint
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。