让3D重建能听懂文字描述,仅凭一张图就能精准还原指定物体。
Ref-SAM3D: Bridging SAM3D with Text for Reference 3D Reconstruction
- 用文本作为高层先验,引导单张图像的3D重建
- 零样本下仅靠语言和2D视图即可生成高质量3D模型
- 适合需要文本控制3D生成的编辑与创作场景
SAM3D在3D物体重建方面表现优异,但其核心局限在于无法根据文本描述重建特定物体,而这对于3D编辑、游戏开发和虚拟环境等实际应用至关重要。为弥补这一空白,我们提出Ref-SAM3D,一种简单有效的SAM3D扩展方法,通过引入文本描述作为高层先验,实现仅凭单张RGB图像与自然语言的文本引导3D重建。大量定性实验表明,Ref-SAM3D在仅依赖自然语言和单一2D视角的情况下,实现了具有竞争力且高保真的零样本重建性能。结果证明,Ref-SAM3D有效弥合了2D视觉线索与3D几何理解之间的鸿沟,为参考引导的3D重建提供了更灵活、易用的新范式。代码已开源:https://github.com/FudanCVL/Ref-SAM3D。
原文摘要 · Abstract (English)
SAM3D has garnered widespread attention for its strong 3D object reconstruction capabilities. However, a key limitation remains: SAM3D cannot reconstruct specific objects referred to by textual descriptions, a capability that is essential for practical applications such as 3D editing, game development, and virtual environments. To address this gap, we introduce Ref-SAM3D, a simple yet effective extension to SAM3D that incorporates textual descriptions as a high-level prior, enabling text-guided 3D reconstruction from a single RGB image. Through extensive qualitative experiments, we show that Ref-SAM3D, guided only by natural language and a single 2D view, delivers competitive and high-fidelity zero-shot reconstruction performance. Our results demonstrate that Ref-SAM3D effectively bridges the gap between 2D visual cues and 3D geometric understanding, offering a more flexible and accessible paradigm for reference-guided 3D reconstruction. Code is available at: https://github.com/FudanCVL/Ref-SAM3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。