用语音节奏特征做说话人识别,发现节奏信息有潜力但受个人表达波动影响。
Rhythm Features for Speaker Identification
- 从语音节奏中提取特征,用深度学习实现无文本说话人识别
- 节奏特征可有效区分说话人,但即兴发言时个体差异大导致效果下降
- 适合研究语音生物特征或非文本识别的学者参考
尽管深度学习模型在说话人识别任务中表现强劲,但主要依赖从频谱图或原始波形中经验性学习的低层音频特征。已有研究表明,个体独特的说话风格显著影响语音信号中语言单元的时间结构(节奏)。这使节奏成为一种强而未被充分重视的说话人身份特征。本文通过深度学习方法,测试了从节奏特征进行无文本说话人识别的可行性。结果表明,节奏信息对说话人识别具有价值,但也发现即兴语境下个体内部变异性较高会削弱其有效性。
原文摘要 · Abstract (English)
While deep learning models have demonstrated robust performance in speaker recognition tasks, they primarily rely on low-level audio features learned empirically from spectrograms or raw waveforms. However, prior work has indicated that idiosyncratic speaking styles heavily influence the temporal structure of linguistic units in speech signals (rhythm). This makes rhythm a strong yet largely overlooked candidate for a speech identity feature. In this paper, we test this hypothesis by applying deep learning methods to perform text-independent speaker identification from rhythm features. Our findings support the usefulness of rhythmic information for speaker recognition tasks but also suggest that high intra-subject variability in ad-hoc speech can degrade its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。