让4D场景编辑更一致且可交互,支持复杂指令自动分解。
4DGS-Craft: Consistent and Interactive 4D Gaussian Splatting Editing
- 引入4D感知的InstructPix2Pix模型,结合几何特征保持视角与时间一致性。
- 通过多视图网格模块迭代优化,提升编辑区域与非编辑区的一致性。
- 基于LLM理解用户意图,自动拆解复杂指令,支持精准交互式编辑。
4D高斯点云渲染(4DGS)编辑近期进展仍面临视角、时间及非编辑区域一致性挑战,且难以处理复杂文本指令。为此,我们提出4DGS-Craft,一个具有一致性与交互性的4DGS编辑框架。首先,引入4D感知的InstructPix2Pix模型,利用初始场景提取的4D VGGT几何特征,捕捉底层4D结构以保障视角与时间一致性。进一步,设计多视图网格模块,在联合优化4D场景的同时迭代精炼多视图输入图像,强化一致性。此外,提出新型高斯选择机制,仅对编辑区域内的高斯进行优化,有效保持非编辑区域稳定性。为提升用户交互能力,构建基于LLM的意图理解模块,通过用户指令模板定义原子操作,并利用大语言模型进行推理,将复杂指令解析为逻辑序列的原子操作,从而实现对复杂命令的高效响应,显著提升编辑可控性与性能。相较现有方法,本框架在一致性与交互性上表现更优。代码将在录用后公开。
原文摘要 · Abstract (English)
Recent advances in 4D Gaussian Splatting (4DGS) editing still face challenges with view, temporal, and non-editing region consistency, as well as with handling complex text instructions. To address these issues, we propose 4DGS-Craft, a consistent and interactive 4DGS editing framework. We first introduce a 4D-aware InstructPix2Pix model to ensure both view and temporal consistency. This model incorporates 4D VGGT geometry features extracted from the initial scene, enabling it to capture underlying 4D geometric structures during editing. We further enhance this model with a multi-view grid module that enforces consistency by iteratively refining multi-view input images while jointly optimizing the underlying 4D scene. Furthermore, we preserve the consistency of non-edited regions through a novel Gaussian selection mechanism, which identifies and optimizes only the Gaussians within the edited regions. Beyond consistency, facilitating user interaction is also crucial for effective 4DGS editing. Therefore, we design an LLM-based module for user intent understanding. This module employs a user instruction template to define atomic editing operations and leverages an LLM for reasoning. As a result, our framework can interpret user intent and decompose complex instructions into a logical sequence of atomic operations, enabling it to handle intricate user commands and further enhance editing performance. Compared to related works, our approach enables more consistent and controllable 4D scene editing. Our code will be made available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。