arXiv:2505.00059cs.CLcs.SD2025-05中稿 · Computer Speech an…被引 8

构建首个包含远距离、情绪化和喊话的语音识别基准数据集

BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition

  • 在98个家庭环境中用手机采集带情绪和喊话的远距离语音
  • 语音识别性能随距离和喊话强度增加而下降,情绪影响识别结果
  • 适合研究真实场景下语音识别与情绪识别的鲁棒性问题

当前自动语音识别(ASR)在多数指标上已接近或达到人类水平,但在真实复杂场景如远距离语音中仍表现不佳。现有挑战多聚焦于距离问题,依赖多麦克风阵列系统。本文提出BERSt数据集,包含98位演员录制的近4小时英语语音,涵盖不同地域及非母语口音,数据通过智能手机在家采集,涉及至少98种声学环境。录音位置达19种,包括遮挡与跨房间情况,且包含7种情绪提示及喊话与正常语速的语音。该数据集公开可用,可用于评估ASR、喊话检测与语音情绪识别(SER)。我们提供了初始的ASR与SER基准,结果显示ASR性能随距离与喊话强度上升而下降,并受情绪类型影响。实验表明BERSt数据集对两类任务均具挑战性,需进一步提升系统在真实场景下的鲁棒性。

原文摘要 · Abstract (English)

Some speech recognition tasks, such as automatic speech recognition (ASR), are approaching or have reached human performance in many reported metrics. Yet, they continue to struggle in complex, real-world, situations, such as with distanced speech. Previous challenges have released datasets to address the issue of distanced ASR, however, the focus remains primarily on distance, specifically relying on multi-microphone array systems. Here we present the B(asic) E(motion) R(andom phrase) S(hou)t(s) (BERSt) dataset. The dataset contains almost 4 hours of English speech from 98 actors with varying regional and non-native accents. The data was collected on smartphones in the actors homes and therefore includes at least 98 different acoustic environments. The data also includes 7 different emotion prompts and both shouted and spoken utterances. The smartphones were places in 19 different positions, including obstructions and being in a different room than the actor. This data is publicly available for use and can be used to evaluate a variety of speech recognition tasks, including: ASR, shout detection, and speech emotion recognition (SER). We provide initial benchmarks for ASR and SER tasks, and find that ASR degrades both with an increase in distance and shout level and shows varied performance depending on the intended emotion. Our results show that the BERSt dataset is challenging for both ASR and SER tasks and continued work is needed to improve the robustness of such systems for more accurate real-world use.

语音识别情绪识别远距离语音真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。