聚焦中文方言语音处理,统一任务与评估标准,推动多方言识别与识别技术发展。
Summary of the ChinaVoices Challenge 2026: Data, Tasks, Baseline, and Methods
- 设立16类方言的识别与语音识别双任务,分受限与开放数据赛道。
- 320小时数据集上,多数系统超越基线,顶级方案在隐藏集保持稳定排名。
- 方言识别依赖声学特征区分,语音识别侧重数据增强与辅助损失优化。
本文总结了中国方言挑战赛2026(ChinaVoices Challenge 2026),旨在建立中文方言语音处理的统一任务定义与评估条件,推动多方言识别与自动语音识别(ASR)的发展。挑战涵盖16种方言类别,设定两个任务:中文多方言识别和中文多方言自动语音识别(ASR)。使用约320小时语音数据,分为参考集、公开评估集和隐藏评估集。两任务共享评估音频,并分别设置受限数据与开放数据赛道。本文介绍任务设置、数据分布、评价指标及Qwen3-ASR-1.7B基线模型,分析排行榜结果与提交系统。共有28支团队提交结果,17支提供系统报告,15支通过合规审查并纳入分析。多数合格系统优于基线,且前三位排名在隐藏评估集上对两个任务均保持不变。方言级结果表明,识别准确率高的类别通常伴随更低的ASR错误率,但两者评估的是相关但独立的能力。领先识别系统普遍利用方言判别性声学表示,而领先ASR系统则强调数据归一化、增强与辅助CTC目标。这些结果为开发与评估中文多方言语音系统提供了实用指导。
原文摘要 · Abstract (English)
This paper summarizes the ChinaVoices Challenge 2026, which aims to establish unified task definitions and evaluation conditions for Chinese dialect speech processing and to advance multi-dialect identification and automatic speech recognition. The challenge covers 16 dialect categories and defines two tasks: Chinese Multi-Dialect Identification and Chinese Multi-Dialect Automatic Speech Recognition (ASR). It uses approximately 320 hours of speech across the Reference Set, Open Evaluation Set, and Hidden Evaluation Set. The two tasks use the same evaluation audio, and each includes restricted-data and open-data tracks. We describe the task settings, data, evaluation metrics, and Qwen3-ASR-1.7B baseline, and analyze the leaderboard results and submitted systems. In total, 28 teams submit results, 17 provide system reports, and systems from 15 teams pass the compliance review and are included in the analysis. Most eligible systems outperform the baseline, and the official top-three order remains unchanged on the Hidden Evaluation Set for both tasks. Dialect-level results show that categories with higher identification accuracy generally have lower ASR error rates, although the tasks assess related but distinct capabilities. Leading identification systems commonly exploit dialect-discriminative acoustic representations, whereas leading ASR systems emphasize data normalization, augmentation, and auxiliary CTC objectives. These results provide practical guidance for developing and evaluating Chinese multi-dialect speech processing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。