构建了覆盖2021东京奥运会的多语言新闻数据集,支持跨语言事件分析。
The 2021 Tokyo Olympics Multilingual News Article Dataset
- 从1918家媒体采集1.09万篇多语言新闻,按赛事事件聚类。
- 涵盖1350个分项赛事,覆盖9种语言、不同语系和文字系统。
- 适合研究跨语言新闻聚类或奥运传播差异的学者使用。
本文介绍了一个涵盖2021年东京奥运会的多语言新闻文章数据集。共收集来自1,918家媒体的10,940篇新闻文章,覆盖1,350个分项赛事,发布时间为2021年7月1日至8月14日。文章涵盖九种不同语言家族、使用不同文字体系的语言。数据通过新闻采集与分析服务获取,经在线聚类算法将报道同一赛事的新闻归为一组,再由人工标注与评估。该数据集旨在为多语言新闻聚类算法提供评测资源,目前此类数据稀缺。也可用于从多视角分析2021年东京奥运会的动态与事件。数据以CSV格式提供,可从CLARIN.SI仓库获取。
原文摘要 · Abstract (English)
In this paper, we introduce a dataset of multilingual news articles covering the 2021 Tokyo Olympics. A total of 10,940 news articles were gathered from 1,918 different publishers, covering 1,350 sub-events of the 2021 Olympics, and published between July 1, 2021, and August 14, 2021. These articles are written in nine languages from different language families and in different scripts. To create the dataset, the raw news articles were first retrieved via a service that collects and analyzes news articles. Then, the articles were grouped using an online clustering algorithm, with each group containing articles reporting on the same sub-event. Finally, the groups were manually annotated and evaluated. The development of this dataset aims to provide a resource for evaluating the performance of multilingual news clustering algorithms, for which limited datasets are available. It can also be used to analyze the dynamics and events of the 2021 Tokyo Olympics from different perspectives. The dataset is available in CSV format and can be accessed from the CLARIN.SI repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。