让语言驱动的2D分割模型学会3D一致性,仅用1160万参数提升3D重建质量
LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency
- 用几何感知的低秩适配层改造2D分割模型,保持3D重建器不变
- 在ScanRefer和Nr3D上,2D掩码与3D重建均显著提升,仅更新1160万参数
- 支持单图或多视角输入,适合需要语言引导3D重建的研究者
文本驱动的3D重建需要能理解自由指令且视角变化下稳定的掩码。我们提出LISA-3D,一种两阶段框架:将指令跟随分割器LISA通过几何感知的低秩适配(LoRA)层进行适应,同时保持SAM-3D重建器冻结。训练时,成对的RGB-D帧与相机位姿定义可微分的重投影损失,强制跨视角一致性,无需额外3D文本标注。部署时,适配后的分割器可从一张RGB图像生成RGBA提示供SAM-3D使用;若有注册的多视角RGB-D数据,可选的对数融合进一步优化提示。在ScanRefer和Nr3D数据集上,几何感知调优同时提升了2D掩码与3D重建效果,仅更新1160万参数。结果清晰分离了几何感知训练增益与可选多视角推理增益,为从语言定位到以物体为中心的3D重建提供模块化路径。
原文摘要 · Abstract (English)
Text-driven 3D reconstruction requires masks that understand free-form instructions and remain stable under viewpoint changes. We present LISA-3D, a two-stage framework that adapts the instruction-following segmenter LISA with geometry-aware Low-Rank Adaptation (LoRA) layers while keeping the SAM-3D reconstructor frozen. During training, paired RGB-D frames and camera poses define a differentiable reprojection loss that enforces cross-view agreement without additional 3D-text annotations. At deployment, the adapted segmenter can produce an RGBA prompt for SAM-3D from one RGB image; when registered RGB-D views are available, optional logit fusion further improves the prompt. On ScanRefer and Nr3D, geometry-aware tuning improves both 2D masks and lifted 3D reconstructions while updating only 11.6M parameters. Our results separate geometry-aware training gains from optional multi-view inference gains, providing a modular route from language grounding to object-centric 3D reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。