arXiv:2502.12583cs.CL2025-02ACL被引 20

用可信合成数据提升大模型长文本推理能力

LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data

  • 通过真实依据和引用提示构建可信推理数据
  • 在多跳推理和LongBench上显著提升模型表现
  • 开源两个数据集,适合长上下文模型训练

尽管长上下文大语言模型发展迅速,但依赖合成数据的数据驱动方法常因缺乏可信性而受限,尤其在长文本推理和问答任务中表现不佳。问题主要源于信息失真、无来源推理及知识冲突。我们提出LongFaith,一种生成可信长文本推理指令数据的新管道。通过融合真实事实与基于引用的推理提示,消除干扰,提升推理链准确性,减少昂贵的验证成本。我们开源了两个合成数据集:LongFaith-SFT 和 LongFaith-PO,系统性地覆盖可验证推理、溯源性和上下文一致性等多维度可信性。在多跳推理数据集和LongBench上的大量实验表明,使用这些数据微调的模型性能显著提升。消融实验证明LongFaith管道具有可扩展性和适应性,适用于多种长上下文大模型开发。

原文摘要 · Abstract (English)

Despite the growing development of long-context large language models (LLMs), data-centric approaches relying on synthetic data have been hindered by issues related to faithfulness, which limit their effectiveness in enhancing model performance on tasks such as long-context reasoning and question answering (QA). These challenges are often exacerbated by misinformation caused by lack of verification, reasoning without attribution, and potential knowledge conflicts. We propose LongFaith, a novel pipeline for synthesizing faithful long-context reasoning instruction datasets. By integrating ground truth and citation-based reasoning prompts, we eliminate distractions and improve the accuracy of reasoning chains, thus mitigating the need for costly verification processes. We open-source two synthesized datasets, LongFaith-SFT and LongFaith-PO, which systematically address multiple dimensions of faithfulness, including verified reasoning, attribution, and contextual grounding. Extensive experiments on multi-hop reasoning datasets and LongBench demonstrate that models fine-tuned on these datasets significantly improve performance. Our ablation studies highlight the scalability and adaptability of the LongFaith pipeline, showcasing its broad applicability in developing long-context LLMs.

长文本推理合成数据可信生成大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。