用自监督方法实现3D心脏CT任意视角自动定位,提升临床诊断效率。
AVP-AP: Self-supervised Automatic View Positioning in 3D cardiac CT via Atlas Prompting
- 基于解剖图谱生成3D标准模板,自监督训练网络映射切片位置。
- 通过刚性配准缩小搜索空间,再用特征相似度优化定位精度。
- 无需人工标注,定位效果优于四名放射科医生,通用性强。
自动视图定位对心脏计算机断层扫描(CT)检查至关重要,但受个体差异和三维搜索空间大影响,难以实现。现有方法依赖繁琐的人工标注,仅能预测固定平面,无法应对任意方向切片在三维体积中的定位需求。本文提出首个基于图谱提示的自监督自动视图定位框架AVP-AP。首先构建3D标准图谱,通过自监督训练网络将切片映射至图谱空间;随后,利用参考CT中对应图谱提示,通过三维图谱与目标CT间的刚性变换确定切片粗略位置,显著缩小搜索范围;最后,基于预训练基础模型的特征空间,最大化预测切片与查询图像的相似性,精修位置。该方法灵活高效,在任意视图定位上平均结构相似性(SSIM)提升19.8%,两腔视图达到9%的SSIM,超越四名放射科医生水平。公开数据集验证了其泛化能力。
原文摘要 · Abstract (English)
Automatic view positioning is crucial for cardiac computed tomography (CT) examinations, including disease diagnosis and surgical planning. However, it is highly challenging due to individual variability and large 3D search space. Existing work needs labor-intensive and time-consuming manual annotations to train view-specific models, which are limited to predicting only a fixed set of planes. However, in real clinical scenarios, the challenge of positioning semantic 2D slices with any orientation into varying coordinate space in arbitrary 3D volume remains unsolved. We thus introduce a novel framework, AVP-AP, the first to use Atlas Prompting for self-supervised Automatic View Positioning in the 3D CT volume. Specifically, this paper first proposes an atlas prompting method, which generates a 3D canonical atlas and trains a network to map slices into their corresponding positions in the atlas space via a self-supervised manner. Then, guided by atlas prompts corresponding to the given query images in a reference CT, we identify the coarse positions of slices in the target CT volume using rigid transformation between the 3D atlas and target CT volume, effectively reducing the search space. Finally, we refine the coarse positions by maximizing the similarity between the predicted slices and the query images in the feature space of a given foundation model. Our framework is flexible and efficient compared to other methods, outperforming other methods by 19.8% average structural similarity (SSIM) in arbitrary view positioning and achieving 9% SSIM in two-chamber view compared to four radiologists. Meanwhile, experiments on a public dataset validate our framework's generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。