首个大规模瑞士德语口语语料库,支持语音识别与方言研究。
SwissGPC v1.0 -- The Swiss German Podcasts Corpus
- 自动化流程构建含5400小时原始音频的语料库
- 保留近5000小时自然对话,覆盖七大方言区
- 适合语音合成、方言识别等真实场景研究
我们提出瑞士GPC v1.0,首个中大规模自发性瑞士德语语音语料库,旨在支持语音识别(ASR)、语音合成(TTS)、方言识别等研究。语料源自瑞士广播电视台及YouTube上的访谈节目和播客,包含约5400小时原始音频。经分割与弱标注后,保留近5000小时语音,覆盖七个主要瑞士德语方言区及标准德语。本文详述语料构建方法,包括自动化标注流程,并提供方言分布、词元数量及分割特征统计。与现有以控制语音为主的语料不同,本语料库捕捉真实自然对话,为实际语音应用提供宝贵资源。
原文摘要 · Abstract (English)
We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to talk shows and podcasts hosted on Schweizer Radio und Fernsehen and YouTube, which contain approximately 5400 hours of raw audio. After segmentation and weak annotation, nearly 5000 hours of speech were retained, covering the seven major Swiss German dialect regions alongside Standard German. We describe the corpus construction methodology, including an automated annotation pipeline, and provide statistics on dialect distribution, token counts, and segmentation characteristics. Unlike existing Swiss German speech corpora, which primarily feature controlled speech, this corpus captures natural, spontaneous conversations, making it a valuable resource for real-world speech applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。