用无标签超声数据预训练模型,提升声速估计的标注效率。
IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data

- 提出基于复数信号的赫米特视觉变压器,直接处理IQ数据。
- 仅需10,000标签即达15.60 m/s误差,标签效率提升3倍以上。
- 适用于超声定量成像,可迁移至不同组织结构的泛化建模。
组织中声速是实现清晰成像和具有诊断价值的前提,但从原始回波通道数据恢复声速本质上是一个非线性逆问题。现有学习方法速度快但依赖大量标注,模拟声速标签成本高,而真实通道数据丰富但无标签。本文提出IQ-JEPA,利用两类数据协同训练。编码器在无标签数据上预训练,通过可见上下文预测被掩码的同相与正交(IQ)区域潜在表示,随后在模拟标签上微调。声速体现在IQ信号中的相位差,对恒定相位偏移不变。所提编码器为直接作用于复数信号的赫米特视觉变压器,其注意力机制对相位及其共轭保持等变,前馈层则对相位不变,从而读取类似传统相干方法的关键量。在79,293次2.5 MHz的Fullwave 2.5仿真下,利用63,435个无标签采集数据预训练,在10,000标签时达到15.60 m/s误差。相比监督训练,标签效率提升约三倍,1,000标签时超过四倍;较InversionNet基线低约2.2倍,全量标签时为8.71 m/s。更多无标签预训练数据可进一步提升性能,表明自监督是关键因素。相同编码器可迁移使用:冻结特征可有效揭示声速与衰减信息,跨层状与腹部体模的分布间预训练仅损失少量精度。本工作被视为定量超声基础模型的重要一步。
原文摘要 · Abstract (English)
The speed of sound in tissue is a prerequisite for well-focused imaging and has diagnostic value, but recovering it from raw pulse-echo channel data is fundamentally a nonlinear inverse problem. Learned solvers are fast yet label hungry. Simulated sound-speed labels are expensive, while abundant real channel data is unlabeled. We propose IQ-JEPA to exploit both data types. An encoder is pretrained without labels to predict the latent representation of masked in-phase and quadrature (IQ) regions from visible context, then fine-tuned on simulated maps. Sound speed appears in the IQ signal as a phase difference, invariant to the constant phase offset. The encoder is a Hermitian vision transformer that operates on the complex signal directly. Its attention is equivariant to that phase and its conjugate-product feed-forward is invariant to it, so the encoder reads a quantity analogous to the one classical coherence methods use. On 79,293 Fullwave 2.5 simulations at 2.5 MHz, pretraining on the 63,435 unlabeled acquisitions reaches 15.60 m/s at 10,000 labels. This is a roughly threefold gain in label efficiency over supervised training, growing to over fourfold at 1,000 labels. It is about 2.2x below an InversionNet baseline, and 8.71 m/s at full labels. The gain still grows with more unlabeled pretraining data. Our comparisons point to self-supervision as the dominant factor. The same encoder transfers. Its frozen features expose sound speed and attenuation, and cross-distribution pretraining between layered and abdominal phantoms costs little accuracy. We see this as a first step toward a foundation model for quantitative ultrasound.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。