arXiv:2512.00677cs.CVcs.AI2025-12被引 3

无需训练即可实现文本驱动的4D场景无缝编辑,保持多视角与时间一致性。

Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer

  • 用多模态扩散变压器结合4D高斯泼溅,实现跨视图与时空融合。
  • 在DyNeRF数据集上编辑保真度优于现有方法,无运动畸变与几何漂移。
  • 适合需要快速、高质量4D内容生成的视觉创作与影视制作人员。

近年来,动态神经辐射场(Dynamic NeRF)和4D高斯泼溅(4DGS)等4D表示取得了进展,实现了动态4D场景重建。然而,由于在编辑过程中难以保证空间与时间维度上的多视角和时序一致性,文本驱动的4D场景编辑仍处于探索阶段。现有方法依赖2D扩散模型独立编辑帧,常导致运动失真、几何漂移和编辑不完整。本文提出Dynamic-eDiTor,一种无需训练的文本驱动4D编辑框架,利用多模态扩散变压器(MM-DiT)与4DGS。该机制包含时空子网格注意力(STGA),用于局部一致的跨视图与时间融合;以及上下文令牌传播(CTP),通过令牌继承与光流引导的令牌替换实现全局传播。二者协同使Dynamic-eDiTor能在不额外训练的前提下,直接优化预训练的4DGS,实现无缝、全局一致的多视角视频编辑。在多视角视频数据集DyNeRF上的大量实验表明,本方法在编辑保真度及多视角与时序一致性方面均优于先前方法。项目页面提供结果与代码:https://di-lee.github.io/dynamic-eDiTor/

原文摘要 · Abstract (English)

Recent progress in 4D representations, such as Dynamic NeRF and 4D Gaussian Splatting (4DGS), has enabled dynamic 4D scene reconstruction. However, text-driven 4D scene editing remains under-explored due to the challenge of ensuring both multi-view and temporal consistency across space and time during editing. Existing studies rely on 2D diffusion models that edit frames independently, often causing motion distortion, geometric drift, and incomplete editing. We introduce Dynamic-eDiTor, a training-free text-driven 4D editing framework leveraging Multimodal Diffusion Transformer (MM-DiT) and 4DGS. This mechanism consists of Spatio-Temporal Sub-Grid Attention (STGA) for locally consistent cross-view and temporal fusion, and Context Token Propagation (CTP) for global propagation via token inheritance and optical-flow-guided token replacement. Together, these components allow Dynamic-eDiTor to perform seamless, globally consistent multi-view video without additional training and directly optimize pre-trained source 4DGS. Extensive experiments on multi-view video dataset DyNeRF demonstrate that our method achieves superior editing fidelity and both multi-view and temporal consistency prior approaches. Project page for results and code: https://di-lee.github.io/dynamic-eDiTor/

4D生成扩散模型文本编辑高斯泼溅

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。