用对话文本替代音频做说话人分离,效果更优。
Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models
- 仅用对话文本进行句级说话人切换检测
- 多模型预测在短对话中表现优于现有音频方法
- 适合需要高精度短对话分离的场景
我们提出一种基于文本的说话人分离新方法,聚焦于对话中的句级说话人切换检测。与依赖音频的系统相比,该方法不依赖音频质量或说话人相似性,仅使用对话转录文本。开发了单预测模型(SPM)和多预测模型(MPM),两者均在短对话中显著提升说话人切换识别能力。基于涵盖多种对话场景的定制数据集,研究发现,文本型SD方法,尤其是MPM,在短对话上下文中表现可媲美最先进的音频系统,且性能更优。本文展示了利用语言特征进行说话人分离的潜力,并强调将语义理解融入系统的必要性,为未来多模态与语义特征驱动的分离研究开辟方向。
原文摘要 · Abstract (English)
We present a novel approach to Speaker Diarization (SD) by leveraging text-based methods focused on Sentence-level Speaker Change Detection within dialogues. Unlike audio-based SD systems, which are often challenged by audio quality and speaker similarity, our approach utilizes the dialogue transcript alone. Two models are developed: the Single Prediction Model (SPM) and the Multiple Prediction Model (MPM), both of which demonstrate significant improvements in identifying speaker changes, particularly in short conversations. Our findings, based on a curated dataset encompassing diverse conversational scenarios, reveal that the text-based SD approach, especially the MPM, performs competitively against state-of-the-art audio-based SD systems, with superior performance in short conversational contexts. This paper not only showcases the potential of leveraging linguistic features for SD but also highlights the importance of integrating semantic understanding into SD systems, opening avenues for future research in multimodal and semantic feature-based diarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。