呼吁将长文本生成纳入大模型研究核心,弥补输出能力短板。
Shifting Long-Context LLMs Research from Input to Output
- 从输入处理转向输出生成,重构长上下文建模思路。
- 长文本生成需兼顾连贯性、逻辑一致与内容丰富度。
- 适合从事创作、规划与复杂推理的AI应用开发者。
近期长上下文大语言模型(LLMs)的研究主要聚焦于处理长输入上下文,显著提升了长文本理解能力。然而,同样关键的长文本生成问题却未得到足够关注。本文倡导自然语言处理研究范式向长输出生成转变。小说创作、长期规划和复杂推理等任务要求模型在理解广泛上下文的同时,生成连贯、富含语境且逻辑一致的长篇内容。这凸显了当前大模型在输出生成方面的明显短板。我们强调这一被忽视领域的价值,并呼吁开发专为高质量长文本生成设计的基础模型,以推动其在真实场景中的广泛应用。
原文摘要 · Abstract (English)
Recent advancements in long-context Large Language Models (LLMs) have primarily concentrated on processing extended input contexts, resulting in significant strides in long-context comprehension. However, the equally critical aspect of generating long-form outputs has received comparatively less attention. This paper advocates for a paradigm shift in NLP research toward addressing the challenges of long-output generation. Tasks such as novel writing, long-term planning, and complex reasoning require models to understand extensive contexts and produce coherent, contextually rich, and logically consistent extended text. These demands highlight a critical gap in current LLM capabilities. We underscore the importance of this under-explored domain and call for focused efforts to develop foundational LLMs tailored for generating high-quality, long-form outputs, which hold immense potential for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。