arXiv:2509.14479cs.SD2025-09被引 4

发布首个长时单人实时语音核磁数据集,支持发音合成与音素识别研究。

A long-form single-speaker real-time MRI speech dataset and benchmark

  • 采集单名美式英语母语者一小时实时口腔动态影像与同步音频。
  • 提供裁剪至声道区域的视频、分句数据、去噪音频等衍生数据。
  • 已建立发音合成与音素识别基线,供后续研究优化参考。

我们发布了包含语音产生过程中实时核磁共振视频与同步音频的南加州大学长时单人(USC Long Single-Speaker, LSS)数据集。该数据集包含一名美式英语母语者约一小时的语音录制,是目前公开可用的最长单人实时语音核磁数据集之一。除了原始的运动学与声学数据,我们还提供了适用于多种下游任务的衍生表示:包括聚焦声道区域的视频、按句子划分的数据、修复与降噪后的音频,以及感兴趣区域的时间序列。我们还在该数据集上对发音合成与音素识别任务进行了基准测试,为未来研究提供了可改进的性能基线。

原文摘要 · Abstract (English)

We release the USC Long Single-Speaker (LSS) dataset containing real-time MRI video of the vocal tract dynamics and simultaneous audio obtained during speech production. This unique dataset contains roughly one hour of video and audio data from a single native speaker of American English, making it one of the longer publicly available single-speaker datasets of real-time MRI speech data. Along with the articulatory and acoustic raw data, we release derived representations of the data that are suitable for a range of downstream tasks. This includes video cropped to the vocal tract region, sentence-level splits of the data, restored and denoised audio, and regions-of-interest timeseries. We also benchmark this dataset on articulatory synthesis and phoneme recognition tasks, providing baseline performance for these tasks on this dataset which future research can aim to improve upon.

语音生成医学影像数据集实时分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。