用结构化推理和强化学习提升大模型金融推理能力
DianJin-R1: Evaluating and Enhancing Financial Reasoning in Large Language Models
- 构建包含财务计算与合规检查的高质量数据集
- 在5个基准上超越非推理模型,单次调用媲美多代理系统
- 适合需要精准金融决策的实操场景
大语言模型在金融领域仍面临有效推理的挑战,任务常需领域知识、精确数值计算及严格遵守合规规则。我们提出DianJin-R1框架,通过推理增强型监督与强化学习解决此问题。核心是DianJin-R1-Data,该数据集源自CFLUE、FinQA及自有合规语料库(中文合规核查,CCC),涵盖多样金融推理场景并附有验证标注。模型DianJin-R1-7B与DianJin-R1-32B基于Qwen2.5-7B-Instruct与Qwen2.5-32B-Instruct,采用结构化格式生成推理步骤与最终答案。为提升推理质量,应用组相对策略优化(GRPO),引入双重奖励信号:鼓励结构化输出,同时奖励答案正确性。在五个基准上评估:三个金融数据集(CFLUE、FinQA、CCC)及两个通用推理基准(MATH-500、GPQA-Diamond)。实验表明,DianJin-R1模型持续优于非推理基线,尤其在复杂金融任务中表现突出。在真实世界CCC数据集上,单次调用推理模型性能达到甚至超过需显著更高算力的多代理系统。结果证明DianJin-R1通过结构化监督与奖励对齐学习,在金融推理方面具有效能,提供可扩展且实用的解决方案。
原文摘要 · Abstract (English)
Effective reasoning remains a core challenge for large language models (LLMs) in the financial domain, where tasks often require domain-specific knowledge, precise numerical calculations, and strict adherence to compliance rules. We propose DianJin-R1, a reasoning-enhanced framework designed to address these challenges through reasoning-augmented supervision and reinforcement learning. Central to our approach is DianJin-R1-Data, a high-quality dataset constructed from CFLUE, FinQA, and a proprietary compliance corpus (Chinese Compliance Check, CCC), combining diverse financial reasoning scenarios with verified annotations. Our models, DianJin-R1-7B and DianJin-R1-32B, are fine-tuned from Qwen2.5-7B-Instruct and Qwen2.5-32B-Instruct using a structured format that generates both reasoning steps and final answers. To further refine reasoning quality, we apply Group Relative Policy Optimization (GRPO), a reinforcement learning method that incorporates dual reward signals: one encouraging structured outputs and another rewarding answer correctness. We evaluate our models on five benchmarks: three financial datasets (CFLUE, FinQA, and CCC) and two general reasoning benchmarks (MATH-500 and GPQA-Diamond). Experimental results show that DianJin-R1 models consistently outperform their non-reasoning counterparts, especially on complex financial tasks. Moreover, on the real-world CCC dataset, our single-call reasoning models match or even surpass the performance of multi-agent systems that require significantly more computational cost. These findings demonstrate the effectiveness of DianJin-R1 in enhancing financial reasoning through structured supervision and reward-aligned learning, offering a scalable and practical solution for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。