用文本和大模型构建越南语语音识别数据集,摆脱对人脸的依赖
VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency

- 通过文本元数据和大模型推理推断说话人身份,不依赖视频中的人脸
- 构建了4715名说话人、约902小时的越南语语音数据集
- 训练模型在鲁棒性和泛化能力上优于现有越南语数据集
语音识别因大规模训练数据集而快速发展,但越南语仍属资源匮乏,现有语料库规模小且声学多样性不足。大多数大规模数据集依赖面部线索将语音与说话人身份关联,限制了仅在有画面时才能采集数据。本文提出一种无脸依赖的数据集构建流程,推出大规模越南语说话人识别数据集VieSpeaker。该方法利用文本元数据和大语言模型推理,从转录文本与上下文信息中推断说话人身份。VieSpeaker包含约902小时语音,来自4,715名说话人。实验表明,基于VieSpeaker训练的模型在鲁棒性与泛化能力上优于现有越南语数据集。本工作验证了无脸依赖数据集构建的可行性,为大规模语音资源建设提供了新方向。
原文摘要 · Abstract (English)
Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial cues to link speech with speaker identities, restricting data collection to recordings where speakers appear on camera. We propose a face-independent dataset construction pipeline and introduce VieSpeaker, a large-scale Vietnamese speaker recognition dataset. Our approach leverages textual metadata and large language model reasoning to infer speaker identities from transcripts and contextual information. VieSpeaker contains approximately 902 hours of speech from 4,715 speakers. Experiments show that models trained on VieSpeaker achieve improved robustness and generalization compared to existing Vietnamese datasets. This work demonstrates the feasibility of face-independent dataset construction and provides a new direction for building large-scale speech resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。