arXiv:2511.03992cs.CV2025-11中稿 · ICME 2026被引 1

让3D高斯点云跨视角一致地理解语言指令,解决视角偏差问题。

Camera-Aware Cross-View Alignment for Referring 3D Gaussian Splatting Segmentation

  • 引入相机条件对齐模块,将相机几何融入点与文本交互
  • 通过高斯级跨视图逻辑对齐,提升多视角预测一致性
  • 适合需要精准跨视角语义定位的研究者和应用

参照3D高斯点云分割(R3DGS)旨在将自由形式的语言查询定位到3D高斯场中。然而,现有方法依赖单视图伪监督,导致视角漂移及跨视图预测不一致。本文提出CaRF(Camera-aware Referring Field),一种面向视图一致性的相机感知跨视图对齐框架。CaRF引入相机条件对齐调制(CAM),将相机几何信息注入高斯-文本交互;并通过高斯级跨视图逻辑对齐(GCLA),在训练中显式对齐同一高斯在标定视图间的指代响应。通过将跨视图差异转化为可优化目标,CaRF实现直接在高斯空间中的几何感知与视图一致推理。在三个基准测试上进行的大量实验表明,CaRF在Ref-LERF、LERF-OVS和3D-OVS上分别提升mIoU 16.8%、4.3%和2.0%,达到当前最优性能。代码已开源:https://github.com/eR3R3/CaRF。

原文摘要 · Abstract (English)

Referring 3D Gaussian Splatting Segmentation (R3DGS) aims to ground free-form language queries in 3D Gaussian fields. However, existing methods rely on single-view pseudo supervision, leading to viewpoint drift and inconsistent predictions across views. We propose CaRF (Camera-aware Referring Field), a camera-aware cross-view alignment framework for view-consistent referring in 3D Gaussian splatting. CaRF introduces Camera-conditioned Alignment Modulation (CAM) to inject camera geometry into Gaussian-text interactions, and Gaussian-level Cross-view Logit Alignment (GCLA) to explicitly align referring responses of the same Gaussians across calibrated views during training. By turning cross-view discrepancy into an optimizable objective, CaRF enables geometry-aware and view-consistent reasoning directly in the Gaussian space. Extensive experiments on three benchmarks demonstrate that CaRF achieves state-of-the-art performance, improving mIoU by 16.8%, 4.3%, and 2.0% on Ref-LERF, LERF-OVS, and 3D-OVS, respectively. Our code is available at https://github.com/eR3R3/CaRF.

3D分割跨视角对齐高斯点云视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。