用跨频协同与动态对齐提升腹部超声视频识别准确率
TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning

- 通过跨频协同适配器增强抗斑点噪声的特征提取
- 多粒度运动感知模块捕捉扫描中的稳定与突变动态
- 动态文本原型生成实现多样扫描条件下的鲁棒对齐
腹部超声对快速无创创伤分诊至关重要,但连续扫查中细微动态线索的解读耗时且依赖操作者。参数高效图像到视频迁移学习(PEIVTL)通过视觉-文本对齐,为超声视频分析提供了新范式。然而,因医生扫描习惯差异带来的时空与语义变化仍严重制约其效果。本文提出TRUST,一种扫描感知的PEIVTL框架,显式建模细粒度时空变异以实现可靠的超声视频理解。首先,提出交叉频率协同适配器(CFCA),在高低频成分间建立互约束,提升重斑点污染下的判别性空间特征提取能力。其次,设计多粒度运动感知模块(MGMA),融合局部时序卷积与运动先验引导的全局自注意力,联合捕捉视图内稳定模式与视图间突变过渡,表征复杂扫描动态。第三,提出视觉查询语义聚合模块(VQSA),根据视觉特征动态生成文本原型,实现对不同扫描条件下类内差异的适应性视觉-文本对齐。在自建腹部创伤超声数据集上的实验表明,TRUST相较最先进方法提升9.63%性能,同时具备更优计算效率。
原文摘要 · Abstract (English)
Abdominal ultrasound is indispensable for rapid, noninvasive trauma triage. However, interpreting the subtle dynamic cues embedded in continuous scanning is time-intensive and operator-dependent. Parameter-Efficient Image-to-Video Transfer Learning (PEIVTL), which efficiently adapts pre-trained image models to the video domain, notably through visual-textual alignment, offers a promising paradigm for ultrasound video analysis. Nevertheless, substantial spatiotemporal and semantic variations arising from physician-dependent scanning practices continue to limit the effectiveness and generalizability of this framework. We propose TRUST, a scan-aware PEIVTL framework that explicitly models fine-grained spatiotemporal variations to enable reliable ultrasound video understanding. First, we introduce a Cross-Frequency Collaborative Adapter (CFCA) that establishes mutual constraints between low- and high-frequency components, enhancing discriminative spatial feature extraction under heavy speckle corruption. Second, we design a Multi-Granularity Motion-Aware (MGMA) module that integrates local temporal convolutions with motion-prior-guided global self-attention, jointly capturing stable intra-view patterns and abrupt inter-view transitions to characterize complex scanning dynamics. Third, a Visual Query Semantic Aggregation (VQSA) module dynamically generates text prototypes conditioned on visual features, enabling adaptive visual-textual alignment robust to intra-class variability under diverse scanning conditions. Experiments on in-house ultrasound trauma datasets demonstrate that TRUST outperforms state-of-the-art methods by 9.63% with superior computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。