arXiv:2507.08660cs.CLcs.LG2025-07Transactions of th…被引 4

自动语音转写虽有错误,但反而更利于说话人识别。

The Impact of Automatic Speech Transcription on Speaker Attribution

  • 用自动语音识别生成的错漏文本进行说话人识别
  • 转写错误率高时识别准确率仍稳定,优于人工转录
  • 适合无音频或音频被匿名化的真实场景

基于语音转写的说话人识别任务旨在通过语言使用模式从转写文本中识别说话人,尤其在音频丢失或被匿名化时极具价值。以往研究多基于人工转录文本,但在真实场景中,通常只有自动语音识别(ASR)系统生成的带错文本。本文首次全面研究了自动转写对说话人识别性能的影响。结果表明,说话人识别对词级转写错误具有惊人鲁棒性,且恢复真实转写的目标与识别性能相关性极低。整体来看,使用错误率更高的ASR转写文本进行识别,效果不劣于甚至优于人工转录,可能因为ASR错误本身包含了反映说话人身份的特定特征。

原文摘要 · Abstract (English)

Speaker attribution from speech transcripts is the task of identifying a speaker from the transcript of their speech based on patterns in their language use. This task is especially useful when the audio is unavailable (e.g. deleted) or unreliable (e.g. anonymized speech). Prior work in this area has primarily focused on the feasibility of attributing speakers using transcripts produced by human annotators. However, in real-world settings, one often only has more errorful transcripts produced by automatic speech recognition (ASR) systems. In this paper, we conduct what is, to our knowledge, the first comprehensive study of the impact of automatic transcription on speaker attribution performance. In particular, we study the extent to which speaker attribution performance degrades in the face of transcription errors, as well as how properties of the ASR system impact attribution. We find that attribution is surprisingly resilient to word-level transcription errors and that the objective of recovering the true transcript is minimally correlated with attribution performance. Overall, our findings suggest that speaker attribution on more errorful transcripts produced by ASR is as good, if not better, than attribution based on human-transcribed data, possibly because ASR transcription errors can capture speaker-specific features revealing of speaker identity.

说话人识别ASR错误转写文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。