arXiv:2511.00260cs.CV2025-11

构建首个用于结肠镜局部配准的基准数据集,解决内窥镜定位难题。

C3VDReg: A Benchmark for Local-to-Local Colonoscopic Registration toward Anatomical Localization

  • 基于真实结肠镜视频与CT重建,生成局部对局部点云配对数据
  • 发现高几何重叠仍无法保证配准准确,重复管状结构是主要瓶颈
  • 适合医学图像配准、智能内窥镜导航研究者参考

解剖感知的结肠镜导航需要将局部内窥镜观测结果定位到稳定的三维参考系中,以支持覆盖范围评估、复检区域识别和CT引导导航。然而,结肠内的刚性点云配准与标准基准存在本质差异:表面局部同质、袋状襞重复、视图高度不完整,且重建深度噪声大。我们提出C3VDReg,一个源自结肠镜3D视频数据集(C3VD)的数据集与基准。每帧通过深度重投影生成源点云,通过匹配相机位姿对CT网格进行射线投射生成目标点云。该基准包含10,015组视角匹配的局部对局部点云对(含2,088组保留测试对),并采用标准化协议评估基线模型:每云点数为8,192,仅源端扰动,固定位姿约定,统一评估指标。关键的是,C3VDReg支持失败模式的系统性分析。我们发现,即使地面真值重叠率达74.3%-93.1%,所有方法的注册召回率依然偏低。通过重叠度、位姿误差及平移分解分析,确定沿重复管状解剖结构的平移歧义是主要瓶颈。这挑战了‘提高重叠度或对应质量即能保证准确配准’的普遍假设,凸显需引入更强的解剖与上下文约束。代码、模型检查点及数据见https://github.com/linzhe001/C3VDReg。

原文摘要 · Abstract (English)

Anatomy-aware colonoscopic navigation requires localizing partial endoscopic observations on a stable 3D reference to support coverage assessment, revisited-region awareness, and CT-guided navigation. However, rigid point cloud registration in the colon differs fundamentally from standard benchmarks: surfaces are locally homogeneous, haustral folds are repetitive, views are highly partial, and reconstructed depth is noisy. We present C3VDReg, a dataset and benchmark derived from the Colonoscopy 3D Video Dataset (C3VD). For each frame, C3VDReg generates source point clouds via depth reprojection and target point clouds by raycasting CT meshes from matched camera poses. The benchmark comprises 10,015 viewpoint-matched partial-to-partial point cloud pairs (including 2,088 held-out test pairs) and evaluates baseline models under a standardized protocol: 8,192 points per cloud, source-only perturbations, fixed pose conventions, and unified metrics. Crucially, C3VDReg enables a systematic investigation of failure modes. We find that high geometric overlap alone is insufficient for reliable pose recovery: despite 74.3-93.1% ground-truth overlap, registration recall remains low across all evaluated methods. Through overlap, pose error, and translation decomposition analyses, we identify translation ambiguity along repetitive tubular anatomy as the primary bottleneck. This challenges the common assumption that increasing overlap or correspondence quality guarantees accurate registration, highlighting the need for stronger anatomical and contextual constraints. Code, model checkpoints, and data are available at https://github.com/linzhe001/C3VDReg .

医学影像点云配准内窥镜导航解剖建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。