arXiv:2606.00890cs.CV2026-06

构建超声视频的群体级神经解剖图谱,实现高效精准标注。

Cohort-Scale Neural Atlases of Ultrasound Video

论文配图:Cohort-Scale Neural Atlases of Ultrasound Video
图 1 · 摘自论文原文
  • 基于DINOv3特征空间,联合训练上千帧数据生成统一模板与每视频嵌入。
  • 在5个数据集上实现单样本/少样本标注迁移,精度媲美强基线,训练仅需分钟级。
  • 嵌入可解释:线性投影揭示群体差异,插值生成合理中间帧,重建缺失画面。

超声是临床最广泛使用的实时成像方式,但逐帧视频标注仍是主要瓶颈:专家标签稀缺且成本高,图像外观受斑点、阴影、衰减及操作者探头位置影响。尤其关键信息常为动态过程,如心脏超声中的左心室运动或肌肉骨骼成像中的肌骨运动。群体解剖图谱可通过将观测数据对齐至共享标准坐标系来分摊标注成本,但现有神经解剖图谱方法多局限于单视频、小规模测试图像集或以物体为中心的图像集合。本文提出一种超声视频的群体规模神经解剖图谱:一个单一标准图谱搭配每视频的生成潜在优化嵌入,于数千帧数据的DINOv3特征空间中联合训练。在五个心脏与肌肉骨骼数据集(含点标记与分割掩码)上,该方法学习到一致的标准化模板,并支持准确的图谱空间标注迁移。在EchoNet-Dynamic和MSK-Bone数据集上,实现了单样本与少样本迁移,精度可比肩强密集对应基线,且在单张消费级GPU上训练仅需数分钟。所学嵌入具可解释性:线性投影揭示结构化群体变化,图像解码器插值生成解剖合理的中间帧,测试时潜在反演可重构被保留帧。结果表明,群体规模神经解剖图谱为降低超声视频分析中的专家标注负担提供了实用且可解释的表示方案。

原文摘要 · Abstract (English)

Ultrasound is the most widely used real-time imaging modality in clinical practice, yet per-frame video annotation remains a major bottleneck: expert labels are scarce and costly, and image appearance varies with speckle, shadowing, attenuation, and operator-dependent probe pose. This is especially limiting because clinically relevant information is often dynamic, from left-ventricular motion in echocardiography to muscle and bone kinematics in musculoskeletal imaging. Population atlases can amortize annotation cost by registering observations to a shared canonical coordinate system, but existing neural atlas methods mainly target single videos, small test-time image sets, or object-centric image collections. We introduce a cohort-scale neural atlas for ultrasound video: a single canonical chart with per-video Generative Latent Optimization embeddings, trained jointly over thousands of frames in DINOv3 feature space. Across five cardiac and musculoskeletal datasets with point landmarks and segmentation masks, our method learns coherent canonical templates and enables accurate atlas-space annotation transfer. On EchoNet-Dynamic and MSK-Bone, it supports single- and few-shot transfer with accuracy competitive with strong dense-correspondence baselines, while training in minutes on a single consumer GPU. The learned embeddings are interpretable: linear projections reveal structured cohort variation, image-decoder interpolation produces anatomically plausible intermediate frames, and test-time latent inversion reconstructs held-out frames through the atlas. These results suggest that cohort-scale neural atlases offer a practical, interpretable representation for reducing expert annotation burden in ultrasound video analysis.

超声视频神经解剖图谱少样本学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。