基于大模型构建超大规模双语时间线数据集,助力事件演化研究
ETimeline: An Extensive Timeline Generation Dataset based on Large Language Model
- 用大模型筛选12万+新闻,生成600个双语时间线
- 覆盖28个新闻领域,含1.3万篇高质量文章
- 适合做事件关系分析、话题生成等任务的研究者
时间线生成对全面理解事件演变至关重要,其目标是将新闻按时间顺序组织,以揭示孤立看待时被遮蔽的模式与趋势,便于追踪故事发展及关键事件间关联。尽管时间线已广泛应用于商业产品,学术研究仍显不足,现有数据集亟需优化。本文提出ETimeline,包含超过13,000篇新闻文章,覆盖600个双语时间线,涵盖28个新闻领域。我们首先构建了超过120,000篇新闻候选池,并利用大语言模型(LLM)管道提升生成性能,最终形成该数据集。数据分析表明其高实用价值。同时提供新闻池数据供后续研究。本工作推动时间线生成研究进展,支持话题生成、事件关系等多样化任务。我们认为该数据集将激发创新研究,弥合学术与产业在技术应用理解间的鸿沟。数据集已公开于https://zenodo.org/records/11392212。
原文摘要 · Abstract (English)
Timeline generation is of great significance for a comprehensive understanding of the development of events over time. Its goal is to organize news chronologically, which helps to identify patterns and trends that may be obscured when viewing news in isolation, making it easier to track the development of stories and understand the interrelationships between key events. Timelines are now common in various commercial products, but academic research in this area is notably scarce. Additionally, the current datasets are in need of refinement for enhanced utility and expanded coverage. In this paper, we propose ETimeline, which encompasses over $13,000$ news articles, spanning $600$ bilingual timelines across $28$ news domains. Specifically, we gather a candidate pool of more than $120,000$ news articles and employ the large language model (LLM) Pipeline to improve performance, ultimately yielding the ETimeline. The data analysis underscores the appeal of ETimeline. Additionally, we also provide the news pool data for further research and analysis. This work contributes to the advancement of timeline generation research and supports a wide range of tasks, including topic generation and event relationships. We believe that this dataset will serve as a catalyst for innovative research and bridge the gap between academia and industry in understanding the practical application of technology services. The dataset is available at https://zenodo.org/records/11392212
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。