arXiv:2502.12714cs.CLcs.SD2025-02NAACL

桌游语音录音成新话者分离挑战,因角色变声导致识别困难。

Playing with Voices: Tabletop Role-Playing Game Recordings as a Diarization Challenge

  • 利用桌游录音构建变声语料库,模拟真实角色扮演中的声音伪装。
  • 两款系统在桌游数据上误判率升高,且威斯佩克严重低估说话人数。
  • 适合研究语音识别鲁棒性与角色化语音处理的学者关注。

本文提出一个概念验证:桌面角色扮演游戏(TTRPG)的音频可作为话者分离系统的挑战。TTRPG以对话为主,参与者常通过变声扮演虚构角色,这种声音转换是沉浸体验的固有特征。这使话者分离系统更难区分真实说话人与角色模仿。我们构建了一个小型TTRPG音频数据集,并与AMI和ICSI语料库进行对比。评估了pyannote.audio和wespeaker两款话者分离系统。结果表明,TTRPG特性导致两类系统误判率上升;尤其wespeaker在TTRPG音频中严重低估说话人数量。研究建议将TTRPG音频作为话者分离系统的新测试基准。

原文摘要 · Abstract (English)

This paper provides a proof of concept that audio of tabletop role-playing games (TTRPG) could serve as a challenge for diarization systems. TTRPGs are carried out mostly by conversation. Participants often alter their voices to indicate that they are talking as a fictional character. Audio processing systems are susceptible to voice conversion with or without technological assistance. TTRPG present a conversational phenomenon in which voice conversion is an inherent characteristic for an immersive gaming experience. This could make it more challenging for diarizers to pick the real speaker and determine that impersonating is just that. We present the creation of a small TTRPG audio dataset and compare it against the AMI and the ICSI corpus. The performance of two diarizers, pyannote.audio and wespeaker, were evaluated. We observed that TTRPGs' properties result in a higher confusion rate for both diarizers. Additionally, wespeaker strongly underestimates the number of speakers in the TTRPG audio files. We propose TTRPG audio as a promising challenge for diarization systems.

语音分离角色扮演音频挑战说话人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。