arXiv:2509.01185cs.CLcs.AI2025-09

用模块化方法生成高质量长文本数据,提升大模型的上下文理解能力。

Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation

  • 通过提示工程与大模型交互生成长文本数据
  • 支持SFT、DPO、GRPO等多种训练目标
  • 适合研究长上下文推理与评估的开发者

大语言模型处理和推理长文本输入的能力对众多实际应用至关重要。然而,该领域进展受限于缺乏高质量、多样且可验证的长上下文数据集,难以用于训练与评估。本文提出一种模块化、可扩展的合成长上下文数据生成框架,基于提示与大模型交互实现。该框架支持监督微调(SFT)、直接偏好优化(DPO)和组相对策略优化(GRPO)等多种训练与对齐目标。包含四种核心生成范式:多轮对话、文档支撑的输入输出对、可验证的指令-响应任务以及长上下文推理示例。通过模板化提示、模型无关架构及带元数据的输出,实现可扩展、可控且目标对齐的数据集构建,助力大模型长上下文能力的发展。

原文摘要 · Abstract (English)

The ability of large language models (LLMs) to process and reason over long textual inputs is critical for a wide range of real-world applications. However, progress in this area is significantly constrained by the absence of high-quality, diverse, and verifiable long-context datasets suitable for both training and evaluation. This work introduces a modular, extensible framework for synthetic long-context data generation via prompt-based interaction with LLMs. The framework supports multiple training and alignment objectives, including Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO). It encompasses four core generation paradigms: multi-turn conversational dialogues, document-grounded input-output pairs, verifiable instruction-response tasks, and long-context reasoning examples. Through templated prompting, a model-agnostic architecture, and metadata-enriched outputs, the proposed approach facilitates scalable, controllable, and purpose-aligned dataset creation for advancing long-context capabilities in LLMs.

长上下文数据生成提示工程LLM训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。