arXiv:2504.14638cs.CV2025-04

用相机位姿插值生成多视角提示,让视觉语言模型更准分割3D物体。

NVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation

  • 基于3D高斯点云,通过相机位姿插值生成多视角硬提示。
  • 无需优化或微调,在多个3D场景中实现更鲁棒的实例分割。
  • 适合希望提升3D目标识别精度且不想训练的新手研究者。

视觉语言模型(VLMs)在图像级视觉感知任务中表现出色的零样本迁移能力,但在需要精确定位和识别单个物体的3D实例分割任务中表现不足。为此,我们提出一种基于3D高斯点云的硬视觉提示方法,利用相机位姿插值在目标物体周围生成多样视角,无需任何2D-3D优化或微调。该方法模拟真实3D视角,通过强制跨视角几何一致性,有效增强现有硬视觉提示。这种免训练策略可无缝集成到已有提示方法中,丰富物体描述特征,使VLM在复杂3D场景中实现更稳健、更准确的实例分割。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have demonstrated impressive zero-shot transfer capabilities in image-level visual perception tasks. However, they fall short in 3D instance-level segmentation tasks that require accurate localization and recognition of individual objects. To bridge this gap, we introduce a novel 3D Gaussian Splatting based hard visual prompting approach that leverages camera interpolation to generate diverse viewpoints around target objects without any 2D-3D optimization or fine-tuning. Our method simulates realistic 3D perspectives, effectively augmenting existing hard visual prompts by enforcing geometric consistency across viewpoints. This training-free strategy seamlessly integrates with prior hard visual prompts, enriching object-descriptive features and enabling VLMs to achieve more robust and accurate 3D instance segmentation in diverse 3D scenes.

3D分割视觉语言模型硬提示高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。