用单目视频实现精准文本编辑4D场景,保持未编辑区域不变。
Mono4DEditor: Text-Driven 4D Scene Editing from Monocular Video via Point-Level Localization of Language-Embedded Gaussians
- 将语言特征嵌入3D高斯点,实现语义定位。
- 两阶段点级定位提升编辑区域精度。
- 适合内容创作与虚拟环境中的动态场景编辑。
基于单目视频重建的4D场景文本驱动编辑是一项具有广泛应用价值但极具挑战性的任务,核心难点在于在复杂动态场景的局部区域实现语义精确编辑,同时保持未编辑内容的完整性。为此,我们提出Mono4DEditor,一种灵活且准确的文本驱动4D场景编辑框架。该方法通过量化CLIP特征增强3D高斯点,构建语言嵌入的动态表示,实现对任意空间区域的高效语义查询。进一步提出两阶段点级定位策略:首先通过CLIP相似性筛选候选高斯点,再精细化调整其空间范围以提升准确性。最后,利用基于扩散的视频编辑模型对定位区域进行定向编辑,并结合光流与涂鸦引导确保空间一致性与时间连贯性。大量实验表明,Mono4DEditor可在多样场景与物体类型上实现高质量文本驱动编辑,同时保留未编辑区域的外观与几何结构,在灵活性与视觉保真度上优于现有方法。
原文摘要 · Abstract (English)
Editing 4D scenes reconstructed from monocular videos based on text prompts is a valuable yet challenging task with broad applications in content creation and virtual environments. The key difficulty lies in achieving semantically precise edits in localized regions of complex, dynamic scenes, while preserving the integrity of unedited content. To address this, we introduce Mono4DEditor, a novel framework for flexible and accurate text-driven 4D scene editing. Our method augments 3D Gaussians with quantized CLIP features to form a language-embedded dynamic representation, enabling efficient semantic querying of arbitrary spatial regions. We further propose a two-stage point-level localization strategy that first selects candidate Gaussians via CLIP similarity and then refines their spatial extent to improve accuracy. Finally, targeted edits are performed on localized regions using a diffusion-based video editing model, with flow and scribble guidance ensuring spatial fidelity and temporal coherence. Extensive experiments demonstrate that Mono4DEditor enables high-quality, text-driven edits across diverse scenes and object types, while preserving the appearance and geometry of unedited areas and surpassing prior approaches in both flexibility and visual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。