arXiv:2502.16684cs.CL2025-02被引 10

用真实用户查询生成大规模多样化的长上下文指令数据

WildLong: Synthesizing Realistic Long-Context Instruction Data at Scale

  • 从真实查询中提取元信息,用图模型建模共现关系,自适应生成数据
  • 在15万条合成数据上微调,长上下文与短上下文任务表现均领先
  • 支持跨文档比较与聚合,适合复杂现实场景的长文本推理

大语言模型(LLMs)扩展上下文窗口后可处理需整合大量信息的任务,但受限于高质量、多样化长上下文指令数据的稀缺。现有数据合成方法多聚焦于事实检索和摘要等单一目标,泛化能力有限。WildLong从真实用户查询中提取元信息,通过图模型建模共现关系,并采用自适应生成策略,实现可扩展的数据合成。其不仅覆盖单文档任务,更支持跨文档比较与聚合等多文档推理。在15万条由WildLong生成的指令-响应对上微调的模型,在多个基准测试中超越现有开源长上下文优化模型,同时在短上下文任务上保持优异性能,无需额外短上下文数据。通过生成更具多样性与真实性的长上下文指令数据集,WildLong提升了模型在长上下文复杂推理中的泛化能力,确立了长上下文数据合成的新范式。

原文摘要 · Abstract (English)

Large language models (LLMs) with extended context windows enable tasks requiring extensive information integration but are limited by the scarcity of high-quality, diverse datasets for long-context instruction tuning. Existing data synthesis methods focus narrowly on objectives like fact retrieval and summarization, restricting their generalizability to complex, real-world tasks. WildLong extracts meta-information from real user queries, models co-occurrence relationships via graph-based methods, and employs adaptive generation to produce scalable data. It extends beyond single-document tasks to support multi-document reasoning, such as cross-document comparison and aggregation. Our models, finetuned on 150K instruction-response pairs synthesized using WildLong, surpasses existing open-source long-context-optimized models across benchmarks while maintaining strong performance on short-context tasks without incorporating supplementary short-context data. By generating a more diverse and realistic long-context instruction dataset, WildLong enhances LLMs' ability to generalize to complex, real-world reasoning over long contexts, establishing a new paradigm for long-context data synthesis.

长上下文数据合成指令微调多文档推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。