构建首个含非言语声音的电影音频分离数据集,解决模型误分笑声尖叫问题。
DnR-nonverbal: Cinematic Audio Source Separation Dataset Containing Non-Verbal Sounds
- 在电影音频中加入笑声、尖叫等非言语声作为语音源
- 实验证明现有模型易将情绪化声音误判为音效
- 适合影视音频处理、声音分离研究者使用
本文提出一个新的电影音频源分离(CASS)数据集DnR-nonverbal,用于处理非言语声音。现有CASS数据集仅包含朗读式语音作为语音源,与真实电影音频不符,导致模型在训练后容易将情绪化声音(如笑声、尖叫)误判为音效而非语音。为此,我们构建了包含笑声、尖叫等非言语声音的语音源的新数据集。实验表明,当前CASS模型存在非言语声音提取偏差问题,而使用本数据集可有效改善合成及真实电影音频中的分离效果。数据集已发布于https://zenodo.org/records/15470640。
原文摘要 · Abstract (English)
We propose a new dataset for cinematic audio source separation (CASS) that handles non-verbal sounds. Existing CASS datasets only contain reading-style sounds as a speech stem. These datasets differ from actual movie audio, which is more likely to include acted-out voices. Consequently, models trained on conventional datasets tend to have issues where emotionally heightened voices, such as laughter and screams, are more easily separated as an effect, not speech. To address this problem, we build a new dataset, DnR-nonverbal. The proposed dataset includes non-verbal sounds like laughter and screams in the speech stem. From the experiments, we reveal the issue of non-verbal sound extraction by the current CASS model and show that our dataset can effectively address the issue in the synthetic and actual movie audio. Our dataset is available at https://zenodo.org/records/15470640.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。