用少量新数据即可高效适配MRI分割模型,提升低资源场景性能。
Role of the Pretraining and the Adaptation data sizes for low-resource real-time MRI video segmentation
- 通过逐步增加预训练数据量,评估模型在新受试者上的泛化能力。
- 仅需15帧新数据即实现98.09%的分割精度,接近全量训练表现。
- 适合语音研究、医疗影像等低资源实时成像场景使用。
实时磁共振成像(rtMRI)常用于语音产生研究,可完整观测发音过程中的喉部结构变化。本研究采用SegNet与UNet模型进行空气-组织边界(ATB)分割任务,通过逐步增加受试者和视频数量进行模型预训练,评估其在两个测试集上的表现:第一组为同源数据中未见受试者与未见视频,相比匹配条件分别提升0.33%(像素分类准确率,PCA)和0.91%(Dice系数);第二组为新数据源中的未见视频,模型达到匹配条件99.63%(PCA)和98.09%(Dice系数)的性能。匹配条件指仅在测试受试者上训练的基准模型。结果表明,微调与适应策略在小样本下仍具显著效果,尤其证实仅需15帧新数据即可实现有效模型适配。
原文摘要 · Abstract (English)
Real-time Magnetic Resonance Imaging (rtMRI) is frequently used in speech production studies as it provides a complete view of the vocal tract during articulation. This study investigates the effectiveness of rtMRI in analyzing vocal tract movements by employing the SegNet and UNet models for Air-Tissue Boundary (ATB)segmentation tasks. We conducted pretraining of a few base models using increasing numbers of subjects and videos, to assess performance on two datasets. First, consisting of unseen subjects with unseen videos from the same data source, achieving 0.33% and 0.91% (Pixel-wise Classification Accuracy (PCA) and Dice Coefficient respectively) better than its matched condition. Second, comprising unseen videos from a new data source, where we obtained an accuracy of 99.63% and 98.09% (PCA and Dice Coefficient respectively) of its matched condition performance. Here, matched condition performance refers to the performance of a model trained only on the test subjects which was set as a benchmark for the other models. Our findings highlight the significance of fine-tuning and adapting models with limited data. Notably, we demonstrated that effective model adaptation can be achieved with as few as 15 rtMRI frames from any new dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。