arXiv:2505.01726cs.CV2025-05ICML被引 3

用概率模型提升3D交互分割的少样本泛化与不确定性判断

Probabilistic Interactive 3D Segmentation with Hierarchical Neural Processes

  • 分层潜在变量捕捉全局场景与物体特征,增强少点击泛化能力
  • 在四个数据集上用更少点击实现更优分割,且能准确预测不确定区域
  • 适合需要高可靠性的3D交互标注场景,如医学或自动驾驶

交互式3D分割通过用户点击生成复杂场景中的精确物体掩码,但仍面临两大挑战:(1)从稀疏点击中有效泛化以生成准确分割;(2)量化预测不确定性以帮助用户识别不可靠区域。本文提出NPISeg3D,一种基于神经过程(Neural Processes, NPs)的概率框架,引入分层潜在变量结构,包含场景特异和物体特异的潜在变量,以同时捕捉全局上下文与物体特定特征,增强少样本泛化能力。此外,设计概率原型调制器,通过物体特异潜在变量自适应调制点击原型,提升模型对物体感知上下文的建模能力并量化预测不确定性。在四个3D点云数据集上的实验表明,NPISeg3D在更少点击下达到更优分割性能,并提供可靠的不确定性估计。

原文摘要 · Abstract (English)

Interactive 3D segmentation has emerged as a promising solution for generating accurate object masks in complex 3D scenes by incorporating user-provided clicks. However, two critical challenges remain underexplored: (1) effectively generalizing from sparse user clicks to produce accurate segmentation, and (2) quantifying predictive uncertainty to help users identify unreliable regions. In this work, we propose NPISeg3D, a novel probabilistic framework that builds upon Neural Processes (NPs) to address these challenges. Specifically, NPISeg3D introduces a hierarchical latent variable structure with scene-specific and object-specific latent variables to enhance few-shot generalization by capturing both global context and object-specific characteristics. Additionally, we design a probabilistic prototype modulator that adaptively modulates click prototypes with object-specific latent variables, improving the model's ability to capture object-aware context and quantify predictive uncertainty. Experiments on four 3D point cloud datasets demonstrate that NPISeg3D achieves superior segmentation performance with fewer clicks while providing reliable uncertainty estimations.

3D分割概率建模少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。