用语言风格分析语音转写文本,实现无音色依赖的说话人识别。
A stylometric analysis of speaker attribution from speech transcripts
- 基于字符、词、句等语言风格特征构建可解释的说话人识别模型
- 标准化转写文本上识别准确率更高,强话题控制下整体表现最优
- 相比黑箱神经模型更易解释,揭示关键区分性语言特征
法医科学家常需在勒索电话、秘密录音、疑似自杀遗书或匿名网络通信等场景中识别未知说话人。传统语音识别依赖声学特征,但当说话人变声或使用语音合成时,该方法失效,仅剩语言内容可用。本文提出将作者身份识别中的语体分析方法应用于语音转写文本,实现说话人归属。我们引入名为StyloSpeaker的语体分析方法,整合字符、词、标记、句子及风格特征,判断两段转写文本是否出自同一说话人。在两种转写格式上评估:一种保留大小写与标点(类书面语),另一种去除这些规范(标准化)。同时控制对话主题程度不同。结果表明,除最强主题控制条件下外,标准化转写文本表现更优;在强主题控制下,整体准确率最高。最后将该可解释模型与黑箱神经方法在相同数据集上对比,分析哪些语言特征最能有效区分说话人。
原文摘要 · Abstract (English)
Forensic scientists often need to identify an unknown speaker or writer in cases such as ransom calls, covert recordings, alleged suicide notes, or anonymous online communications, among many others. Speaker recognition in the speech domain usually examines phonetic or acoustic properties of a voice, and these methods can be accurate and robust under certain conditions. However, if a speaker disguises their voice or employs text-to-speech software, vocal properties may no longer be reliable, leaving only their linguistic content available for analysis. Authorship attribution methods traditionally use syntactic, semantic, and related linguistic information to identify writers of written text (authorship attribution). In this paper, we apply a content-based authorship approach to speech that has been transcribed into text, using what a speaker says to attribute speech to individuals (speaker attribution). We introduce a stylometric method, StyloSpeaker, which incorporates character, word, token, sentence, and style features from the stylometric literature on authorship, to assess whether two transcripts were produced by the same speaker. We evaluate this method on two types of transcript formatting: one approximating prescriptive written text with capitalization and punctuation and another normalized style that removes these conventions. The transcripts' conversation topics are also controlled to varying degrees. We find generally higher attribution performance on normalized transcripts, except under the strongest topic control condition, in which overall performance is highest. Finally, we compare this more explainable stylometric model to black-box neural approaches on the same data and investigate which stylistic features most effectively distinguish speakers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。