arXiv:2512.24314cs.CL2025-12

金融大模型分阶段训练,提升推理与智能决策能力

QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs

  • 分四阶段训练:预训练→金融微调→推理强化→智能体强化
  • 在权威金融评测中表现超越现有模型,推理与智能体能力显著提升
  • 适合金融行业落地应用,为工业级大模型提供可复用范式

金融领域大模型的专用增强长期是产业应用重点。尽管以往模型如BloombergGPT和Baichuan-Finance主要聚焦知识增强,但金融服务复杂性的加深,促使业界对兼具领域知识、强金融推理及智能体能力的模型需求日益增长。本文提出金融大模型QianfanHuijin,并构建一种通用的多阶段训练范式。方法从金融语料持续预训练(CPT)开始,建立知识基础;随后通过逐步细化的后训练流程:金融SFT、金融推理强化学习(RL)、金融智能体强化学习(RL),最终对接真实业务场景的通用强化学习。实证结果表明,QianfanHuijin在多个权威金融基准上表现优异。消融实验进一步验证,推理强化与智能体强化阶段分别带来显著能力提升。该研究支持了渐进式后训练的设计动机,表明此精细分阶段方法有望成为各类工业增强大模型的主流范式。

原文摘要 · Abstract (English)

Domain-specific enhancement of Large Language Models (LLMs) within the financial context has long been a focal point of industrial application. While previous models such as BloombergGPT and Baichuan-Finance primarily focused on knowledge enhancement, the deepening complexity of financial services has driven a growing demand for models that possess not only domain knowledge but also robust financial reasoning and agentic capabilities. In this paper, we present QianfanHuijin, a financial domain LLM, and propose a generalizable multi-stage training paradigm for industrial model enhancement. Our approach begins with Continual Pre-training (CPT) on financial corpora to consolidate the knowledge base. This is followed by a fine-grained Post-training pipeline designed with increasing specificity: starting with Financial SFT, progressing to Finance Reasoning RL and Finance Agentic RL, and culminating in General RL aligned with real-world business scenarios. Empirical results demonstrate that QianfanHuijin achieves superior performance across various authoritative financial benchmarks. Furthermore, ablation studies confirm that the targeted Reasoning RL and Agentic RL stages yield significant gains in their respective capabilities. These findings validate our motivation and suggest that this fine-grained, progressive post-training methodology is poised to become a mainstream paradigm for various industrial-enhanced LLMs.

金融大模型多阶段训练推理能力智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。