arXiv:2503.00751cs.CLcs.AI2025-03ACL被引 31

Rapid通过写作规划与信息发现,高效生成连贯长文本。

RAPID: Efficient Retrieval-Augmented Long Text Generation with Writing Planning and Information Discovery

  • 先生成带检索的初步提纲,减少幻觉
  • 约束搜索提升信息获取效率,降低延迟
  • 计划引导生成,保持主题一致性,适合写百科类内容

生成知识密集型、全面的长篇文本(如百科全书文章)仍是大语言模型的重大挑战,需精确整合事实并维持全文主题连贯性。现有方法如直接生成或多智能体讨论常出现幻觉、主题不连贯和高延迟问题。为此,我们提出RAPID,一种高效的检索增强型长文本生成框架。该框架包含三个模块:(1) 基于检索的初步提纲生成,以减少幻觉;(2) 属性约束搜索,实现高效信息发现;(3) 计划引导的文章生成,增强连贯性。在新构建的基准数据集FreshWiki-2024上的大量实验表明,RAPID在多项评估指标(如长文本生成质量、提纲质量、延迟等)上显著优于现有最先进方法。本工作为自动化长文本生成提供了鲁棒且高效的解决方案。

原文摘要 · Abstract (English)

Generating knowledge-intensive and comprehensive long texts, such as encyclopedia articles, remains significant challenges for Large Language Models. It requires not only the precise integration of facts but also the maintenance of thematic coherence throughout the article. Existing methods, such as direct generation and multi-agent discussion, often struggle with issues like hallucinations, topic incoherence, and significant latency. To address these challenges, we propose RAPID, an efficient retrieval-augmented long text generation framework. RAPID consists of three main modules: (1) Retrieval-augmented preliminary outline generation to reduce hallucinations, (2) Attribute-constrained search for efficient information discovery, (3) Plan-guided article generation for enhanced coherence. Extensive experiments on our newly compiled benchmark dataset, FreshWiki-2024, demonstrate that RAPID significantly outperforms state-of-the-art methods across a wide range of evaluation metrics (e.g. long-text generation, outline quality, latency, etc). Our work provides a robust and efficient solution to the challenges of automated long-text generation.

长文本生成检索增强写作规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。