arXiv:2603.05094cs.SD2026-03

构建台湾方言音文数据集,提升本地语音模型表现

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling

  • 用验证-生成-评审流程筛选并扩增58万条音文指令对
  • 在声学基准测试中准确率达49.1%,比零样本基线高6.5%
  • 适合做方言语音语言模型研究或本地化智能助手开发

大型音频-语言模型(LALMs)常因缺乏专门语料而难以处理地方性口音。我们提出TW-Sound580K,一个基于验证-生成-评审(VGC)协议构建的台湾地区音文指令数据集。该流程利用双语音识别(Dual-ASR)验证过滤出52.2万条原始音频片段,并通过教师模型扩展为58万条高质量音文对。其有效性通过Tai-LALM验证:该模型以DeSTA 2.5-Audio初始化主干网络,并引入动态双ASR仲裁策略优化推理时的转录选择。在TAU基准测试中,Tai-LALM达到49.1%准确率,相比零样本基线(使用ASR文本条件时为42.6%)提升6.5个百分点,证明结合区域性语料与严格筛选、动态仲裁机制可显著提升模型在本地语音任务上的性能。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) typically struggle with localized dialectal prosody due to the scarcity of specialized corpora. We present TW-Sound580K, a Taiwanese audio-text instruction dataset developed through a Verify-Generate-Critique (VGC) protocol. This pipeline leverages Dual-ASR validation to filter 522K raw clips, subsequently expanding them into 580,000 high-fidelity instruction pairs using a teacher model. The dataset's utility is demonstrated through Tai-LALM, which fine-tunes a DeSTA 2.5-Audio-initialized backbone and incorporates a dynamic Dual-ASR Arbitration strategy to optimize transcription selection during inference. On the TAU Benchmark, Tai-LALM reaches 49.1% accuracy, marking a 6.5% absolute improvement over the zero-shot baseline (42.6% with ASR text conditioning). This confirms that integrating regional corpora with rigorous curation and dynamic arbitration significantly enhances LALM performance on localized speech.

音文对方言建模数据集ASR仲裁

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。