arXiv:2607.26249cs.CLcs.CY2026-07

构建了美国宗教广播的海量转录语料库,支持传播研究与语音分析。

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

  • 从785个直播流抓取15分钟片段,覆盖2000+电台,生成超70万条录音。
  • 每条录音经自动化流程转录并标注说话人,共6000万+已分段语音行。
  • 适合研究宗教媒体内容、社会议题讨论及小众领域语音处理的学者。

宗教广播是美国广泛但研究不足的大众传播形式,其内容分析长期受限于缺乏大规模转录数据。本文介绍一个语料库,包含2025年7月通过网络直播流捕获的英文宗教广播转录内容。在为期一个月的时间内,以滚动方式从785个不同频道录制15分钟片段,这些频道共同重播了超过两千个调频/调幅电台信号,共产生70万余条录音,超过6000万条已进行说话人分离的语音行。所有录音均通过自动化流程完成转录和说话人区分,并利用大语言模型对节目格式与主题进行分割与标注。语料库以流元数据、录制元数据和转录行的关联表形式组织,可用于跨区域与教派的宗教广播描述性研究,分析宗教媒体中社会政治议题的讨论方式,以及在未充分代表领域的语音处理研究。

原文摘要 · Abstract (English)

Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained by the absence of large-scale transcript data. This Data Descriptor presents a corpus of transcribed English-language religious radio broadcasts captured from live webstreams over a one-month period in July 2025. Fifteen-minute segments were recorded on a rolling schedule from 785 distinct streams, which together rebroadcast the signals of more than two thousand AM and FM stations, yielding over 700,000 recordings and more than 60 million diarized lines of speech. Each recording was transcribed and speaker-diarized with an automated pipeline, and segmented and labeled by programming format and topic using a large language model. The corpus is organized as linked tables of stream metadata, recording metadata, and transcript lines. It supports descriptive study of religious broadcasting across regions and traditions, analysis of how social and political issues are discussed in religious media, and speech-processing research in an underrepresented domain.

语料库宗教媒体语音分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。