arXiv:2506.14427eess.AScs.MM2025-06被引 3

构建多模态多场景多语言语音分离数据集,解决标注难与模型泛化差问题。

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

  • 融合音视频信息自动生成伪标签,高效构建大规模数据集。
  • M3SD涵盖真实网络视频,覆盖多种语言、场景与说话人组合。
  • 开源数据与代码,助力语音分离技术研究与模型验证。

在语音分离领域,技术发展受限于数据资源不足与深度学习模型泛化能力差两大问题。为此,我们提出一种自动化构建语音分离数据集的方法,通过结合音频与视频信息,生成大规模数据的更准确伪标签。基于该方法,我们发布了多模态、多场景、多语言语音分离数据集(M3SD)。该数据集源自真实网络视频,具有高度多样性。相关数据与代码已开源至 https://huggingface.co/spaces/OldDragon/m3sd。

原文摘要 · Abstract (English)

In the field of speaker diarization, the development of technology is constrained by two problems: insufficient data resources and poor generalization ability of deep learning models. To address these two problems, firstly, we propose an automated method for constructing speaker diarization datasets, which generates more accurate pseudo-labels for massive data through the combination of audio and video. Relying on this method, we have released Multi-modal, Multi-scenario and Multi-language Speaker Diarization (M3SD) datasets. This dataset is derived from real network videos and is highly diverse. Our dataset and code have been open-sourced at https://huggingface.co/spaces/OldDragon/m3sd.

语音分离多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。