arXiv:2410.15511cs.IR2024-10EMNLP被引 4

用树状结构检索让长文本生成更深入全面

ConTReGen: Context-driven Tree-structured Retrieval for Open-domain Long-form Text Generation

  • 基于上下文驱动的树状检索,分层挖掘查询细节
  • 在LFQA、ODSUM等数据集上超越现有最优模型
  • 适合需要深度知识融合的复杂问答场景

开放域长文本生成需对复杂问题提供兼具广度与深度的连贯回答。现有迭代式检索增强生成方法难以深入挖掘复杂查询的各个维度,且难以有效整合多源知识。本文提出ConTReGen框架,采用上下文驱动的树状检索策略,结合自顶向下的分层深入探索与自底向上的系统性整合,确保多维度信息的完整覆盖与连贯融合。在多个数据集(包括LFQA、ODSUM)及新构建的ODSUM-WikiHow上进行的大量实验表明,ConTReGen显著优于现有最先进RAG模型。

原文摘要 · Abstract (English)

Open-domain long-form text generation requires generating coherent, comprehensive responses that address complex queries with both breadth and depth. This task is challenging due to the need to accurately capture diverse facets of input queries. Existing iterative retrieval-augmented generation (RAG) approaches often struggle to delve deeply into each facet of complex queries and integrate knowledge from various sources effectively. This paper introduces ConTReGen, a novel framework that employs a context-driven, tree-structured retrieval approach to enhance the depth and relevance of retrieved content. ConTReGen integrates a hierarchical, top-down in-depth exploration of query facets with a systematic bottom-up synthesis, ensuring comprehensive coverage and coherent integration of multifaceted information. Extensive experiments on multiple datasets, including LFQA and ODSUM, alongside a newly introduced dataset, ODSUM-WikiHow, demonstrate that ConTReGen outperforms existing state-of-the-art RAG models.

长文本生成检索增强树状结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。