让大模型用可验证的预测动作实现金融推理与时间序列预测的统一。
Reasoning through Verifiable Forecast Actions: Consistency-Grounded RL for Financial LLMs

- 通过结构化预测动作连接语言推理与时间序列建模,提升可解释性。
- 在10年大规模基准上,8B模型推理准确率提升25.9%。
- 适合需要可验证、数值化决策的金融场景应用。
金融市场具有高度非平稳性、信噪比低,且严重依赖新闻、公司基本面和宏观经济信号等外部信息。现有方法或抽象时间序列成文本,或割裂预测与语言推理,造成定性推理与定量结果脱节。为此,我们提出StockR1,一种增强时间序列的大型语言模型,通过可验证的预测动作统一股票预测与金融推理。基于工具调用设计,模型先输出结构化、可解释的市场展望动作,再以此为条件调用时间序列解码器生成未来分布轨迹,从而支持更精准的问题回答与金融推理。整个流程通过强化学习优化,奖励同时考量答案有效性、预测准确性及动作与实际时间序列动态的一致性;同时引入样本级不确定性标量重加权,使模型能适应市场动态中的不确定性变化。在为期10年的大规模基准上评估,StockR1持续优于时间序列基线与通用大模型,在4B和8B规模下分别提升推理准确率17.7%与25.9%。结果表明,结构化预测动作能有效促进语言推理与时间预测的协同,使大模型能够进行可验证、可解释且数值化的决策推理。
原文摘要 · Abstract (English)
Financial markets are characterized by extreme non-stationarity, low signal-to-noise ratios, and strong dependence on external information such as news, company fundamentals, and macroeconomic signals. Yet, existing approaches either abstract time-series into text or decouple forecasting from language-based reasoning, leading to a fundamental mismatch between qualitative reasoning and quantitative outcomes. To address this, we introduce StockR1, a time-series-enhanced LLM that unifies stock forecasting and financial reasoning through a verifiable forecast action. Based on a tool-call design, the model first emits a forecast action, which is a structured and interpretable representation of its qualitative market outlook. It then invokes a time-series decoder conditioned on this action to generate distributional future trajectories, leading to more informed question answering and financial reasoning. We optimize the full pipeline with reinforcement learning, where rewards jointly reflect answer validity, forecast accuracy, and consistency between generated actions and observed time-series dynamics. In addition, rewards are reweighted by a sample-level uncertainty scalar, encouraging the model to accommodate varying uncertainty in market dynamics. We evaluate StockR1 on financial question answering and stock forecasting over a large-scale 10-year benchmark. Our method consistently outperforms time-series baselines and general-purpose LLMs, improving reasoning accuracy by 17.7% (4B) and 25.9% (8B). These findings demonstrate that structuring the forecast actions establishes a powerful synergy between language reasoning and temporal prediction, enabling LLMs to reason through verifiable, interpretable, and numerically grounded decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。