arXiv:2601.03948cs.AIq-fin.TR2026-01被引 1

用过程验证解决金融强化学习中的奖励噪声问题

Trade-R1: Bridging Verifiable Rewards to Stochastic Environments via Process-Level Reasoning Verification

  • 将长文本推理转为RAG结构化任务,通过三元一致性检验过滤噪声
  • 动态语义奖励使跨市场泛化能力更强,推理一致性最高
  • 适合研究金融AI决策、可信强化学习的从业者

强化学习(RL)已使大语言模型在数学与编程等可验证奖励领域展现卓越推理能力。然而,将其扩展至金融市场面临挑战:尽管回报可验证,但市场具有随机性,导致标准强化学习陷入奖励劫持。为此,我们提出Trade-R1,一种通过过程级推理验证将可验证奖励引入随机环境的训练框架。核心创新是将长篇金融文档推理评估转化为结构化的检索增强生成(RAG)任务,并构建三角一致性度量,评估检索证据、推理链与决策之间的两两对齐,作为噪声市场回报的有效性过滤器。我们探索两种奖励融合策略:固定效应语义奖励(FSR)以稳定对齐信号,动态效应语义奖励(DSR)实现耦合幅度优化。在多国资产选择任务上的实验表明,该范式有效减少奖励劫持,其中DSR在跨市场泛化上表现更优,同时保持最高推理一致性。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has enabled Large Language Models (LLMs) to achieve remarkable reasoning in domains like mathematics and coding, where verifiable rewards provide clear signals. However, extending this paradigm to financial decision is challenged by the market's stochastic nature: rewards are verifiable but inherently noisy, causing standard RL to degenerate into reward hacking. To address this, we propose Trade-R1, a model training framework that bridges verifiable rewards to stochastic environments via process-level reasoning verification. Our key innovation is a verification method that transforms the problem of evaluating reasoning over lengthy financial documents into a structured Retrieval-Augmented Generation (RAG) task. We construct a triangular consistency metric, assessing pairwise alignment between retrieved evidence, reasoning chains, and decisions to serve as a validity filter for noisy market returns. We explore two reward integration strategies: Fixed-effect Semantic Reward (FSR) for stable alignment signals, and Dynamic-effect Semantic Reward (DSR) for coupled magnitude optimization. Experiments on different country asset selection demonstrate that our paradigm reduces reward hacking, with DSR achieving superior cross-market generalization while maintaining the highest reasoning consistency.

强化学习金融AI推理验证RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。