用自提问迭代生成新闻时间线,突破信息过载瓶颈。
Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization
- 通过不断自问自答,动态优化新闻事件关联与摘要
- 在开放域时间线总结任务中表现超越封闭域模型
- 专为真实新闻场景设计,适合媒体与情报分析
在信息快速变化的背景下,从海量事件相关文本中构建连贯时间线变得愈发重要且困难。挑战在于将相关文档聚合,围绕核心主题构建有意义的事件图谱。本文提出CHRONOS——基于迭代自提问的因果标题检索方法,用于开放域新闻时间线总结。该方法利用大语言模型(LLMs)在每轮中反思事件关联性,并针对特定新闻话题提出新问题,从在线或离线知识库中获取信息,持续生成并更新时间线摘要。此外,我们构建了Open-TLS数据集,包含由专业记者撰写的近期新闻主题时间线,用于评估开放域时间线总结任务,该任务因信息过载导致无法从网络中找到全面的相关文档。实验表明,CHRONOS不仅擅长开放域时间线总结,其性能甚至可媲美为封闭域任务设计的现有最先进系统。
原文摘要 · Abstract (English)
In the fast-changing realm of information, the capacity to construct coherent timelines from extensive event-related content has become increasingly significant and challenging. The complexity arises in aggregating related documents to build a meaningful event graph around a central topic. This paper proposes CHRONOS - Causal Headline Retrieval for Open-domain News Timeline SummarizatiOn via Iterative Self-Questioning, which offers a fresh perspective on the integration of Large Language Models (LLMs) to tackle the task of Timeline Summarization (TLS). By iteratively reflecting on how events are linked and posing new questions regarding a specific news topic to gather information online or from an offline knowledge base, LLMs produce and refresh chronological summaries based on documents retrieved in each round. Furthermore, we curate Open-TLS, a novel dataset of timelines on recent news topics authored by professional journalists to evaluate open-domain TLS where information overload makes it impossible to find comprehensive relevant documents from the web. Our experiments indicate that CHRONOS is not only adept at open-domain timeline summarization, but it also rivals the performance of existing state-of-the-art systems designed for closed-domain applications, where a related news corpus is provided for summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。