利用语音特征提升对话主题分割的鲁棒性
Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features
- 联合训练文本与语音编码器,捕捉句间音频线索
- 在多语言数据集上显著优于纯文本模型
- 对语音识别错误更敏感,适合嘈杂环境
口语内容如在线视频和播客常涵盖多个主题,自动主题分割对用户导航和下游应用至关重要。然而,现有方法未充分挖掘声学特征,仍有提升空间。本文提出一种多模态方法,同时微调文本编码器与孪生语音编码器,捕捉句间声学线索。在大规模YouTube视频数据集上的实验表明,该模型显著优于纯文本及多模态基线。此外,模型对ASR噪声更具鲁棒性,在葡萄牙语、德语和英语三个附加数据集上表现优于更大规模的纯文本基线,凸显了学习到的声学特征在鲁棒主题分割中的价值。
原文摘要 · Abstract (English)
Spoken content, such as online videos and podcasts, often spans multiple topics, which makes automatic topic segmentation essential for user navigation and downstream applications. However, current methods do not fully leverage acoustic features, leaving room for improvement. We propose a multi-modal approach that fine-tunes both a text encoder and a Siamese audio encoder, capturing acoustic cues around sentence boundaries. Experiments on a large-scale dataset of YouTube videos show substantial gains over text-only and multi-modal baselines. Our model also proves more resilient to ASR noise and outperforms a larger text-only baseline on three additional datasets in Portuguese, German, and English, underscoring the value of learned acoustic features for robust topic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。