arXiv:2509.15568cs.CLcs.AI2025-09AAAI被引 4

用结构化主题和多智能体辩论高效生成长文本训练数据

LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs

  • 按书籍分类体系组织主题,用多智能体辩论生成多样化话题
  • 通过轻量BM25检索相关文档,拼接成128K-token样本
  • 降低计算与工程成本,适合长文本训练研究者使用

高质量长上下文数据对训练可处理长文档的大语言模型至关重要,但现有基于相关性聚合的合成方法存在计算效率低的问题。我们提出LiteLong,一种通过结构化主题组织与多智能体辩论实现资源高效的长上下文数据合成方法。该方法利用BISAC图书分类体系提供全面的层级主题结构,并通过多个大模型的辩论机制在该结构内生成多样且高质量的主题。针对每个主题,采用轻量级BM25检索相关文档,并拼接成128K token的训练样本。在HELMET和Ruler基准上的实验表明,LiteLong达到具有竞争力的长上下文性能,且可无缝集成其他长依赖增强方法。LiteLong通过降低计算与数据工程成本,使高质量长上下文数据合成更具可及性,推动长上下文语言建模研究发展。

原文摘要 · Abstract (English)

High-quality long-context data is essential for training large language models (LLMs) capable of processing extensive documents, yet existing synthesis approaches using relevance-based aggregation face challenges of computational efficiency. We present LiteLong, a resource-efficient method for synthesizing long-context data through structured topic organization and multi-agent debate. Our approach leverages the BISAC book classification system to provide a comprehensive hierarchical topic organization, and then employs a debate mechanism with multiple LLMs to generate diverse, high-quality topics within this structure. For each topic, we use lightweight BM25 retrieval to obtain relevant documents and concatenate them into 128K-token training samples. Experiments on HELMET and Ruler benchmarks demonstrate that LiteLong achieves competitive long-context performance and can seamlessly integrate with other long-dependency enhancement methods. LiteLong makes high-quality long-context data synthesis more accessible by reducing both computational and data engineering costs, facilitating further research in long-context language training.

长文本生成数据合成多智能体资源效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。