让3D场景编辑更真实:几何与外观分离,支持灵活修改且不破坏结构。
TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking

- 用网格引导生成对齐的3D锚点,实现几何与外观解耦编辑
- 单卡RTX Pro 6000上每项编辑约10分钟,8个场景6类任务均超越现有方法
- 适合需要高保真3D场景修改的研究者和内容创作者
现有3D高斯泼溅(3DGS)编辑方法主要关注外观修改,难以灵活调整几何结构的同时保持结构完整性和场景一致性。为此,我们提出TRACE,一种基于网格引导的3DGS编辑框架,可自动将显式3D几何与高斯场景对齐,并解耦几何锚定与外观调和。首先,多视角3D锚点合成(在我们的MV-TRACE数据集上训练)生成几何对齐的编辑锚点;其次,可穿戴几何对齐(TGA)完成粗到精的网格-场景配准。然后,情境视频掩码(CVM)将投影的3D锚点融入自回归视频扩散流水线,同步调和其外观并保持多视角一致性。我们在六个编辑类别下的八个独立场景上评估TRACE,结果显示每项编辑仅需约10分钟(单张NVIDIA RTX Pro 6000 GPU),实验表明其在编辑多样性、结构完整性、语义对齐、多视角一致性和视觉质量方面全面优于现有方法。
原文摘要 · Abstract (English)
Existing 3D Gaussian Splatting (3DGS) editing methods primarily focus on appearance modification and often struggle to support flexible geometry editing while preserving structural integrity and scene-consistent appearance. To address this limitation, we present TRACE, a mesh-guided 3DGS editing framework that automatically aligns explicit 3D geometry with Gaussian scenes and decouples Geometric Anchoring from Appearance Harmonization. First, Multi-view 3D-Anchor Synthesis, trained on our MV-TRACE dataset for scene-coherent object addition and modification, generates geometrically aligned editing anchors, while Tangible Geometry Alignment (TGA) performs coarse-to-fine mesh-scene registration. Then, Contextual Video Masking (CVM) integrates projected 3D anchors into an autoregressive video diffusion pipeline, harmonizing their appearance with the surrounding scene while maintaining multi-view consistency. We evaluate TRACE on eight held-out scenes across six editing categories. TRACE completes each edit in approximately 10 minutes on a single NVIDIA RTX Pro 6000 GPU. Extensive experiments demonstrate consistent improvements over existing methods in editing versatility, structural integrity, semantic alignment, multi-view consistency, and visual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。