arXiv:2505.19429cs.CL2025-05被引 1

构建1.3万集播客数据集,助力精准识别精彩片段

Rhapsody: A Dataset for Highlight Detection in Podcasts

  • 用YouTube播放回放频次标注播客片段,构建段落级高亮标签
  • 微调模型在长音频中表现显著优于零样本大模型
  • 适合研究语音内容理解与个性化推荐的学者和工程师

播客已成全球五亿用户日常伴侣。面对海量内容,精彩片段能帮助听众快速把握节目核心并决定是否完整收听。然而,由于内容结构松散且时长较长,自动识别亮点极具挑战。本文提出Rhapsody数据集,包含13,000个播客节目,每段均配有基于YouTube‘最常回放’功能生成的高亮评分。我们将播客高亮检测建模为段落级二分类任务。实验对比了零样本提示语言模型与轻量级微调模型的表现。结果表明,即使采用GPT-4o和Gemini等先进模型,其零样本性能仍有限;而使用领域内数据微调后的模型显著超越基线,尤其在结合语音信号特征与文本转录的基础上效果更优。研究揭示了长篇口语媒体中细粒度信息获取的深层挑战。

原文摘要 · Abstract (English)

Podcasts have become daily companions for half a billion users. Given the enormous amount of podcast content available, highlights provide a valuable signal that helps viewers get the gist of an episode and decide if they want to invest in listening to it in its entirety. However, identifying highlights automatically is challenging due to the unstructured and long-form nature of the content. We introduce Rhapsody, a dataset of 13K podcast episodes paired with segment-level highlight scores derived from YouTube's 'most replayed' feature. We frame the podcast highlight detection as a segment-level binary classification task. We explore various baseline approaches, including zero-shot prompting of language models and lightweight fine-tuned language models using segment-level classification heads. Our experimental results indicate that even state-of-the-art language models like GPT-4o and Gemini struggle with this task, while models fine-tuned with in-domain data significantly outperform their zero-shot performance. The fine-tuned model benefits from leveraging both speech signal features and transcripts. These findings highlight the challenges for fine-grained information access in long-form spoken media.

播客分析高亮检测语音理解数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。