首个面向开放域时间线总结的大型模型,提升新闻演变理解能力。
TIM: A Large-Scale Dataset and large Timeline Intelligence Model for Open-domain Timeline Summarization
- 构建1000+主题、3000+标注实例的大规模数据集
- 通过渐进式优化实现更准确的时间线摘要
- 适合关注新闻演化分析与时间推理的研究者
开放域时间线总结(TLS)对监控新闻话题演变至关重要。现有方法多依赖通用大语言模型从检索到的新闻中提取相关时间点并进行总结,虽具备零样本摘要和时间定位能力,但在判断话题相关性及理解话题演变方面表现不足,常导致内容冗余或时间戳错误。为此,我们提出首个开放域时间线智能模型TIM,具备高效生成开放域时间线总结的能力。首先,构建一个大规模TLS数据集,包含超过1000个新闻话题和3000多个标注的TLS实例。其次,提出渐进式优化策略:先通过指令微调提升摘要与去噪能力,再引入新颖的双对齐奖励学习方法,融合语义与时间双重视角,增强对话题演化规律的理解。实验表明,该策略显著提升了时间线总结性能,在开放域场景下验证了模型有效性。
原文摘要 · Abstract (English)
Open-domain Timeline Summarization (TLS) is crucial for monitoring the evolution of news topics. To identify changes in news topics, existing methods typically employ general Large Language Models (LLMs) to summarize relevant timestamps from retrieved news. While general LLMs demonstrate capabilities in zero-shot news summarization and timestamp localization, they struggle with assessing topic relevance and understanding topic evolution. Consequently, the summarized information often includes irrelevant details or inaccurate timestamps. To address these issues, we propose the first large Timeline Intelligence Model (TIM) for open-domain TLS, which is capable of effectively summarizing open-domain timelines. Specifically, we begin by presenting a large-scale TLS dataset, comprising over 1,000 news topics and more than 3,000 annotated TLS instances. Furthermore, we propose a progressive optimization strategy, which gradually enhance summarization performance. It employs instruction tuning to enhance summarization and topic-irrelevant information filtering capabilities. Following this, it exploits a novel dual-alignment reward learning method that incorporates both semantic and temporal perspectives, thereby improving the understanding of topic evolution principles. Through this progressive optimization strategy, TIM demonstrates a robust ability to summarize open-domain timelines. Extensive experiments in open-domain demonstrate the effectiveness of our TIM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。