arXiv:2410.06519cs.CL2024-10EMNLP被引 4

让小模型也能高效处理长文本,靠结构化笔记和过滤机制。

SEGMENT+: Long Text Processing with Short-Context Language Models

  • 用结构化笔记和过滤模块管理信息流,控制输入内容。
  • 在长文档问答和找针入海任务中显著提升性能。
  • 适合资源有限但需处理长文本的场景,如嵌入式系统。

当前对扩展语言模型输入容量的需求日益增长。然而,单纯增大上下文窗口并不能保证在各类长文本任务中的鲁棒表现,例如理解长文档或从冗长嘈杂的数据中提取细节信息。为此,我们提出SEGMENT+,一种通用框架,使语言模型能在有限上下文窗口内高效处理长输入。该框架利用结构化笔记与过滤模块控制信息流动,实现可调控且可解释的处理过程。我们在多种模型规模下进行了广泛实验,聚焦于长文档问答与找针入海任务,结果表明SEGMENT+能有效提升性能。

原文摘要 · Abstract (English)

There is a growing interest in expanding the input capacity of language models (LMs) across various domains. However, simply increasing the context window does not guarantee robust performance across diverse long-input processing tasks, such as understanding extensive documents and extracting detailed information from lengthy and noisy data. In response, we introduce SEGMENT+, a general framework that enables LMs to handle extended inputs within limited context windows efficiently. SEGMENT+ utilizes structured notes and a filtering module to manage information flow, resulting in a system that is both controllable and interpretable. Our extensive experiments across various model sizes, focusing on long-document question-answering and Needle-in-a-Haystack tasks, demonstrate the effectiveness of SEGMENT+ in improving performance.

长文本处理小模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。