构建首个大规模越南语口语对话数据集,助力聊天机器人与客服系统语音识别。
VietSuperSpeech: A Large-Scale Vietnamese Conversational Speech Dataset for ASR Fine-Tuning in Chatbot, Customer Support, and Call Center Applications
- 从YouTube真实对话视频采集52,023段口语音频,覆盖日常交流场景。
- 数据集含267.39小时语音,经筛选后用于训练和测试,平均每个语句266字符。
- 填补越南语语音识别在非正式口语场景的空白,适合对话系统开发者使用。
我们推出VietSuperSpeech,一个包含52,023个音视频对、总计267.39小时的大型越南语自动语音识别(ASR)数据集,专注于自然口语对话。不同于现有以朗读、新闻播报或有声书为主的越南语语料库,该数据集源自四个公开的YouTube频道,涵盖日常生活对话、个人博客、海外越侨交流及非正式评论,真实反映聊天机器人、客服中心和热线服务中的语言风格。所有音频统一为16 kHz单声道PCM WAV格式,并切分为3-30秒的语句。转录通过基于Zipformer-30M-RNNT-6000h模型(Nguyen, 2025)的Sherpa-ONNX工具进行伪标注,该模型在6,000小时越南语语音上预训练。经质量过滤后,数据集划分为46,822个训练样本(240.67小时)和5,201个开发/测试样本(26.72小时),采用固定随机种子。每条文本平均266字符,共包含1380万条带声调的越南语字符。实验表明,尽管已有VLSP2020、VIET_BUD500、VietSpeech等数据集覆盖正式与朗读语音,但均未专门针对口语化、自发性语言,而VietSuperSpeech正填补这一关键空白。数据集已公开发布于https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech。
原文摘要 · Abstract (English)
We introduce VietSuperSpeech, a large-scale Vietnamese automatic speech recognition (ASR) dataset of 52,023 audio-text pairs totaling 267.39 hours, with a distinctive focus on casual conversational speech. Unlike existing Vietnamese ASR corpora that predominantly feature read speech, news narration, or audiobook content, VietSuperSpeech is sourced from four publicly accessible YouTube channels spanning everyday conversation, personal vlogging, overseas Vietnamese community dialogue, and informal commentary - the very speech styles encountered in real-world chatbot, customer support, call center, and hotline deployments. All audio is standardized to 16 kHz mono PCM WAV and segmented into 3-30 second utterances. Transcriptions are generated via pseudo-labeling using the Zipformer-30M-RNNT-6000h model (Nguyen, 2025) deployed through Sherpa-ONNX, pre-trained on 6,000 hours of Vietnamese speech. After quality filtering, the dataset is split into 46,822 training samples (240.67 hours) and 5,201 development/test samples (26.72 hours) with a fixed random seed. The text averages 266 characters per utterance, totaling 13.8 million fully diacritically marked Vietnamese characters. We demonstrate that VietSuperSpeech fills a critical gap in the Vietnamese ASR ecosystem: while corpora such as VLSP2020, VIET_BUD500, VietSpeech, FLEURS, VietMed, Sub-GigaSpeech2-Vi, viVoice, and Sub-PhoAudioBook provide broad coverage of formal and read speech, none specifically targets the casual, spontaneous register indispensable for conversational AI applications. VietSuperSpeech is publicly released at https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。