构建首个百万级英文播客数据集,支持大规模播客生态研究。
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus
- 收集2020年5-6月公开RSS源的110万集播客转录文本。
- 包含37万集的音频特征与说话人轮次,110万集的说话人角色标注。
- 适合对播客内容、结构和听众响应感兴趣的学者与研究者。
播客通过独特的按需模式为海量听众提供多样化内容,但受限于数据稀缺,难以开展大规模计算分析。为此,我们构建了一个涵盖超过1.1M集英文播客转录文本的大型数据集,覆盖2020年5月至6月期间所有可通过公开RSS源获取的英语播客。该数据不仅包含文本,还包含37万集的音频特征与说话人轮次,以及全部1.1M集的说话人角色推断与其他元数据。基于此数据,我们对播客生态的内容、结构及响应特性进行了基础性探索。本数据集与分析为该流行且具有影响力媒介的持续计算研究开辟了道路。
原文摘要 · Abstract (English)
Podcasts provide highly diverse content to a massive listener base through a unique on-demand modality. However, limited data has prevented large-scale computational analysis of the podcast ecosystem. To fill this gap, we introduce a massive dataset of over 1.1M podcast transcripts that is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020. This data is not limited to text, but rather includes audio features and speaker turns for a subset of 370K episodes, and speaker role inferences and other metadata for all 1.1M episodes. Using this data, we also conduct a foundational investigation into the content, structure, and responsiveness of this ecosystem. Together, our data and analyses open the door to continued computational research of this popular and impactful medium.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。