arXiv:2603.01353cs.LGcs.AI2026-03

用合成数据提升领域大模型的推理能力,日本金融领域验证有效。

Constructing Synthetic Instruction Datasets for Improving Reasoning in Domain-Specific LLMs: A Case Study in the Japanese Financial Domain

  • 从领域词汇出发生成带思维链的合成指令数据
  • 构建95亿token金融领域数据集,显著提升模型表现
  • 开源数据与模型,适合垂直领域研究者使用

在将大语言模型适配特定领域时,同时具备领域专业知识和推理能力仍是迫切挑战。本文提出一种通用方法,可为任意领域构建高质量合成指令数据,起点为领域专用词汇。以金融领域为例,我们构建了一个包含约95亿token的大型指令数据集,其中包含链式思维(Chain-of-Thought)推理轨迹。评估结果表明,该数据集在金融基准测试中优于基线模型,验证了方法的有效性。同时报告了推理轨迹长度对性能的影响及其局限性。最后,我们已将模型与数据集公开于https://huggingface.co/nri-ai。

原文摘要 · Abstract (English)

In adapting LLMs to specific domains, achieving both domain expertise and reasoning ability remains an urgent challenge. This study proposes a general method for constructing high-quality synthetic instruction data for any domain, starting from domain-specific vocabulary. As a demonstration, we applied this method to the financial domain and constructed a large-scale instruction dataset totaling approximately 9.5 billion tokens with Chain-of-Thought reasoning traces. Evaluation results confirmed performance improvements over baseline models on financial benchmarks, demonstrating the effectiveness of our approach. We also report findings on the impact of reasoning trace length on performance and its limitations. Lastly, we open-source our models and datasets on https://huggingface.co/nri-ai .

合成数据领域模型推理增强金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。