用语音信息辅助快速核磁成像,提升发音动态可视化质量
SIREM: Speech-Informed MRI Reconstruction with Learned Sampling

- 融合语音与磁共振数据,通过声学预测发音结构
- 在螺旋采样模式下实现更高吞吐量,重建结果更符合解剖结构
- 适合语音科学和临床评估中需要高速动态成像的研究者
实时磁共振成像(rtMRI)可非侵入式观察发音过程中的动态声道运动,对语音科学和临床评估具有重要价值。然而,其受限于空间分辨率、时间分辨率与采集速度之间的权衡,常导致k空间欠采样,影响重建质量。本文提出SIREM框架,利用同步语音作为跨模态先验:发音时声道形态与声音特征相关,部分图像内容可从音频预测。SIREM将每帧建模为音频驱动成分与MRI驱动成分的融合,通过空间权重图实现。音频分支从语音中预测发音体结构,MRI分支则基于采集的k空间数据重构互补内容。进一步引入可学习的螺旋臂软权重策略,实现对采样方式与多模态融合交互的可微分研究。该方法统一了音频预测、图像重建与采样优化。在USC语音rtMRI基准上评估显示,相比传统插值、小波压缩感知与总变差等基线方法,SIREM在显著更高的吞吐量下仍保持解剖合理的声道结构。该工作建立了多模态语音引导rtMRI重建的初始基准,揭示了同步语音作为辅助先验在快速重建中的潜力。代码已开源。
原文摘要 · Abstract (English)
Real-time magnetic resonance imaging (rtMRI) of speech production enables non-invasive visualization of dynamic vocal-tract motion and is valuable for speech science and clinical assessment. However, rtMRI is fundamentally constrained by trade-offs among spatial resolution, temporal resolution, and acquisition speed, often leading to undersampled k-space measurements and degraded reconstructions. We propose SIREM, a speech-informed MRI reconstruction framework that uses synchronized speech as a cross-modal prior. The central idea is that vocal-tract configurations during speech are correlated with the produced acoustics, making part of the image content predictable from audio. SIREM models each frame as a fusion of an audio-driven component and an MRI-driven component through a spatial weighting map. The audio branch predicts articulator-related structure from speech, while the MRI branch reconstructs complementary content from measured k-space data. We further introduce a learnable soft weighting profile over spiral arms, enabling a differentiable study of how k-space arm usage interacts with speech-informed fusion. This yields a unified multimodal formulation that combines audio-driven prediction, MRI reconstruction, and sampling adaptation. We evaluate SIREM on the USC speech rtMRI benchmark against standard baselines, including gridding, wavelet-based compressed sensing, and total variation. SIREM introduces a speech-informed reconstruction paradigm that operates in a substantially higher-throughput regime than iterative methods while preserving anatomically plausible vocal-tract structure. These results establish an initial benchmark for multimodal speech-informed rtMRI reconstruction and highlight the potential of synchronized speech as an auxiliary prior for fast reconstruction. The source code is available at https://github.com/mdhasanai/SIREM
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。