构建5.1小时俄语语音数据集,自动标注韵律特征提升语音合成质量。
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
- 基于语义语音活动检测与多语音识别集成,实现上下文保留的精准分段
- 构建5.1小时多源俄语语料库,支持音调、重音、音标等丰富标注
- 适用于语音降噪与语音合成任务,适合关注俄语语音处理的研究者
我们提出Balalaika,一个开源的数据驱动型语音处理流水线,用于音频处理与韵律感知标注。该流程结合语义语音活动检测(VAD)实现上下文保留的分段,采用多语音识别器集成与ROVER共识解码,并可选择性保留词级时间戳,随后进行自动质量与说话人纯净度过滤。文本进一步增强标点恢复、词汇重音与"\textipa{e}/\textipa{He}"归一化及国际音标(IPA)音素标注。使用Balalaika,我们构建了5.1小时多源俄语语料库,包含丰富标注。在同等训练预算下,该语料库在语音降噪与语音合成任务中均表现出一致提升;消融实验验证了重音与标点的互补收益,且更严格的MOS过滤可进一步提升合成质量。数据集已公开于HuggingFace。
原文摘要 · Abstract (English)
We introduce Balalaika, an open-source, data-centric pipeline for processing audio and producing prosody-aware annotations. It combines semantic VAD for context-preserving segmentation, multi-ASR ensembling with ROVER consensus decoding, while retaining optional word-level timestamps, followed by automatic quality and speaker-purity filtering. The text is further enriched with punctuation restoration, lexical stress and "\textipa{e}/\textipa{He}" normalization, and IPA phonemes. Using Balalaika, we build a 5.1k-hour multi-source Russian corpus with rich annotations, and show consistent gains under equalized training budgets for both speech denoising and TTS; ablations confirm complementary benefits of stress and punctuation and improved synthesis with stricter MOS filtering. The datasets are publicly available at \href{https://huggingface.co/collections/lab260/balalaika-dataset}{\underline{\textbf{HuggingFace}}}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。