无需训练和原始相机,单参考图即可高精度提取3D高斯场景中的目标物体。
Seed2GS: Camera-Free, Training-Free Object Extraction from 3D Gaussian Scenes via a Single Reference-View Grounding

- 用一个参考视图分离目标身份与三维覆盖,避免重复检测。
- 在LERF-MASK上达到92.1% mIoU,推理仅需9.3秒。
- 适合无原始相机或无法训练的3D场景交互编辑任务。
从预构建的3D高斯溅射(3DGS)场景中提取目标物体可支持交互式3D编辑。现有方法要么每场景需训练数十分钟,牺牲精度,或依赖原始重建相机,而预构建资产常不包含这些信息。本文提出Seed2GS,无需原始相机或场景特定表示训练,即达到最高报告的LERF-MASK准确率。其核心思想是将目标身份与3D覆盖解耦:QD-SAM3从多个开集候选中选取一个可靠参考掩码并固定身份;随后通过种子提升与自适应虚拟轨道暴露新视角下的物体,同时跟踪传播种子,无需重复检测。由于场景保持冻结,每个高斯仅受临时前景logit监督。在LERF-MASK上,Seed2GS实现92.1%的平均交并比(mIoU),计算延迟仅9.3秒,较最强的场景训练基线高3.7个百分点,较最接近的无相机基线高7.6个百分点。每场景使用一个固定测试参考时,完整流程仍保持91.1% mIoU;将预测种子替换为真实掩码仅提升0.72个百分点。在3D-OVS上,达到95.7% mIoU。
原文摘要 · Abstract (English)
Extracting a target object from a pre-built 3D Gaussian Splatting (3DGS) scene enables interactive 3D editing. Existing methods either train for tens of minutes per scene, sacrifice accuracy, or require original reconstruction cameras that pre-built assets may not include. We present Seed2GS, which achieves the highest reported LERF-MASK accuracy without original reconstruction cameras or scene-specific representation training. Its key insight is to separate target identity from 3D coverage. QD-SAM3 selects one reliable reference mask from several open-vocabulary candidates, fixing identity once. Seed lift and visibility-adaptive virtual orbits then expose the object from new viewpoints, while tracking propagates the seed without repeated detection. Because the scene remains frozen, these masks supervise only one temporary foreground logit per Gaussian. On LERF-MASK, Seed2GS reaches 92.1% mean intersection over union (mIoU) with a measured compute-only latency of 9.3 seconds, 3.7 points above the strongest scene-trained baseline and 7.6 points above the closest camera-free baseline. With one fixed test reference per scene, the complete pipeline retains 91.1% mIoU; replacing its predicted seed with a ground-truth mask improves mIoU by only 0.72 points. On 3D-OVS, Seed2GS reaches 95.7% mIoU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。