arXiv:2608.11787cs.CLcs.AI2026-08

用强化学习训练金融建议模型,效果超越商用大模型。

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

  • 用分组相对策略优化提升语言模型的金融建议能力。
  • 在因果审计下,利润提升达商用模型的两倍(0.0228 vs 0.0104)。
  • 结合判官独立评估,发现传统评价可能遗漏真实商业价值。

从商业记录生成可操作的财务建议需融合数值推理、领域知识与合理判断,同时避免损害业务的建议。直接监督困难:历史决策未必最优,高质量自由格式标签成本高昂。本文将财务建议生成建模为强化学习问题,使用分组相对策略优化(GRPO)微调开源语言模型。奖励函数基于大模型作为裁判的评分体系,从多个二元维度评估建议质量,并加入安全门控防止有害推荐。由于仅依赖大模型评估无法确认改进是否反映真实商业价值,而非适应裁判偏好,因此引入独立于裁判的因果审计——标准双重稳健条件平均处理效应(CATE)估计器。在观测性离策略审计中,训练后的模型估计的毛利润提升约为最强商业基线的两倍(0.0228 对 0.0104),且下行率最低、负尾部风险最小。值得注意的是,两种评估对基线的排名不一致:未训练的基础模型在裁判评分中垫底,但在因果审计中位列第二,表明审计捕捉到裁判未察觉的信号。结果表明,基于金融语境的奖励信号与GRPO结合,能生成显著优于商用大模型的商业建议,且判官独立的因果审计是比验证更有力的评估工具。

原文摘要 · Abstract (English)

Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision is difficult: historical decisions are not necessarily optimal, and high-quality free-form labels are expensive to obtain. We formulate financial advice generation as a reinforcement learning problem and fine-tune an open-weight language model using Group Relative Policy Optimization (GRPO). Our reward is an LLM-as-a-judge rubric that scores each recommendation across multiple binary dimensions of advice quality, augmented with a safety gate for harm prevention. Since LLM-based evaluation alone cannot confirm whether improvements reflect genuine business value rather than adaptation to the judge, we complement it with a judge-independent audit based on a standard doubly-robust Conditional Average Treatment Effect (CATE) estimator. Under this observational off-policy audit, our trained LLM achieves approximately twice the estimated gross-profit lift of the strongest evaluated commercial baseline ($0.0228$ vs.\ $0.0104$), together with the lowest downside rate and the least negative tail risk of any policy evaluated. Notably, the two evaluations do not rank the baselines identically: the untrained base model places last on the judge rubric but second on the causal audit, indicating that the audit captures a signal the judge does not. Our results demonstrate that GRPO with a finance-grounded reward signal can produce substantially more useful business recommendations than commercial LLMs, and that a judge-independent causal audit is a valuable complement to, rather than a confirmation of, LLM-as-a-judge assessment in financial NLP.

金融AI强化学习大模型评估因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。