arXiv:2501.06229cs.CVcs.SD2025-01被引 5

开源3D MRI声道数据集,用于深度学习自动分割。

Open-Source Manually Annotated Vocal Tract Database for Automatic Segmentation from 3D MRI Using Deep Learning: Benchmarking 2D and 3D Convolutional and Transformer Networks

  • 构建开源手动标注的3D MRI声道数据集
  • 对比2D/3D卷积与Transformer网络性能
  • 为语音研究提供高质量标注基准

从磁共振成像(MRI)数据中准确分割声道对语音和发音应用至关重要。手动分割耗时且易出错。本研究旨在评估深度学习算法在3D MRI上自动分割声道的有效性。为此,我们构建了一个开源的手动标注3D MRI声道数据集,并在此基础上对2D与3D卷积神经网络及Transformer网络进行了基准测试。该数据集包含多例高分辨率3D MRI扫描,覆盖不同个体与发音状态,为自动分割模型训练与评估提供了可靠基准。实验结果表明,3D卷积与Transformer架构在复杂形状建模上表现更优,尤其在边界精度与一致性方面优于传统2D方法。该工作为语音生成、发音建模等任务提供了可复现的数据与方法基础。

原文摘要 · Abstract (English)

Accurate segmentation of the vocal tract from magnetic resonance imaging (MRI) data is essential for various voice and speech applications. Manual segmentation is time intensive and susceptible to errors. This study aimed to evaluate the efficacy of deep learning algorithms for automatic vocal tract segmentation from 3D MRI.

医学图像深度学习声道分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。