构建了3000万条多语言评论数据集,支持跨领域研究。
YT-30M: A multi-lingual multi-category dataset of YouTube comments
- 从YouTube抓取3223万条评论,覆盖多个频道类别。
- 提供10万条随机样本(YT-100K)用于高效实验。
- 适合做多语言文本分析与社交媒体研究的人使用。
本文介绍了两个大规模多语言评论数据集——YT-30M(及子集YT-100K),均来自YouTube。研究基于较小的样本集YT-100K进行分析。这两个数据集均已公开:YT-30M(完整版)包含32,236,173条评论,而YT-100K为从中随机抽取的108,694条评论。每条评论包含视频ID、评论ID、评论者姓名、评论者频道ID、评论内容、点赞数、原始频道ID及所属频道类别(如'News & Politics'、'Science & Technology'等)。数据可用于多语言、多领域社交文本研究。
原文摘要 · Abstract (English)
This paper introduces two large-scale multilingual comment datasets, YT-30M (and YT-100K) from YouTube. The analysis in this paper is performed on a smaller sample (YT-100K) of YT-30M. Both the datasets: YT-30M (full) and YT-100K (randomly selected 100K sample from YT-30M) are publicly released for further research. YT-30M (YT-100K) contains 32236173 (108694) comments posted by YouTube channel that belong to YouTube categories. Each comment is associated with a video ID, comment ID, commentor name, commentor channel ID, comment text, upvotes, original channel ID and category of the YouTube channel (e.g., 'News & Politics', 'Science & Technology', etc.).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。