arXiv:2603.07223cs.LG2026-03被引 2

通过高质量数据蒸馏与难度感知训练,提升金融大模型准确性与可靠性。

Unlocking Data Value in Finance: A Study on Distillation and Difficulty-Aware Training

  • 采用多阶段知识蒸馏生成高可信链式思考数据
  • 在9个金融任务上超越同规模开源模型表现
  • 适合金融AI研究者与需要高精度推理的从业者

大型语言模型虽具强大泛化能力,但在金融领域部署仍面临专业术语密集、数值推理要求严格及容错率极低等挑战。本研究通过受控实证表明,特定垂直领域的性能主要取决于后训练数据的质量及其难度与可验证性特征。我们构建了两个数据集:ODA-Fin-SFT-318k,通过多阶段蒸馏与验证生成高质量链式思考监督信号;ODA-Fin-RL-12k,精选难但可验证的任务以平衡奖励精度与任务多样性。基于标准SFT与强化学习流程,高质量链式思考蒸馏为SFT奠定坚实基础,而难度与可验证性感知采样显著提升强化学习泛化能力。在涵盖通用金融任务、情感分析与数值推理的九个基准上,我们的ODA-Fin-RL-8B持续优于同规模开源金融大模型。我们公开发布ODA-Fin-SFT-318k与ODA-Fin-RL-12k数据集及训练模型,推动以数据为中心的金融AI研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated strong general capabilities, yet their deployment in finance remains challenging due to dense domain-specific terminology, stringent numerical reasoning requirements, and low tolerance for factual errors. We conduct a controlled empirical study showing that in specialized vertical domains, performance is largely determined by the quality and difficulty/verifiability profile of post-training data. We introduce \textbf{ODA-Fin-SFT-318k}, constructed via multi-stage distillation and verification to produce high-quality Chain-of-Thought supervision, and \textbf{ODA-Fin-RL-12k}, curated for hard-but-verifiable tasks that balance reward precision and task diversity. Using standard SFT and RL pipelines, we show that high-quality CoT distillation establishes a robust foundation during SFT, while difficulty- and verifiability-aware sampling improves RL generalization. Evaluated on nine benchmarks spanning general financial tasks, sentiment analysis, and numerical reasoning, our ODA-Fin-RL-8B consistently surpasses open-source state-of-the-art (SOTA) financial LLMs of comparable size. We release our ODA-Fin-SFT-318k and ODA-Fin-RL-12k datasets, along with trained models to advance data-centric financial AI research.

金融AI知识蒸馏链式思考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。