arXiv:2605.18466cs.CV2026-05

利用语音信息提升实时核磁共振声道分割精度

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI

论文配图:Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI
图 1 · 摘自论文原文
  • 用语音生成位置先验,引导声道结构定位
  • 跨模态对比预训练实现音视频特征对齐
  • 仅需影像即可推理,适合临床实时应用

在实时核磁共振(rtMRI)中分割声道运动部件是低对比度、快速运动与空间分辨率受限的动态图像分割难题。尽管rtMRI可同步获取语音信号,现有方法却忽略此信息;少数多模态方法在无音频时无法使用。本文提出三阶段框架:训练时利用语音与音位监督,推理时仅需rtMRI图像;将音位表示转为解剖位置先验,通过双层跨模态对比预训练对齐视觉与音频编码器,并以交叉注意力解码器融合特征,实现多模态知识向单模态推理的迁移。在75-Speaker~Annot-16和USC-TIMIT数据集上验证,该方法优于现有单模态与多模态方法,证明多模态监督能有效提升分割精度与临床实用性。

原文摘要 · Abstract (English)

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisitions may provide synchronized acoustic signals, existing methods discard this information, and the few multimodal approaches that incorporate audio cannot be deployed when audio is unavailable. We propose a three-stage framework that leverages acoustic and phonological supervision during training while requiring only the rtMRI image at inference: phonological representations are converted into spatial bounding-box priors for articulator localization, visual and acoustic encoders are aligned via dual-level cross-modal contrastive pretraining, and the learned representations are fused through a cross-attention decoder, effectively transferring multimodal knowledge into a single-modality inference pipeline. Evaluated on 75-Speaker~Annot-16 and USC-TIMIT datasets, our method outperforms existing unimodal and multimodal methods, demonstrating that multimodal supervision provides transferable benefits for precise and clinically deployable vocal tract segmentation.

声学分割跨模态学习实时医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。