通过局部隐式特征解耦,实现大尺度与细粒度动态场景的多视角建模。
LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature Decoupling
- 将动态场景分解为种子定义的局部空间,分区域捕捉运动
- 分离静态与动态特征,结合生成时序高斯点建模运动
- 首次实现复杂大尺度动态场景重建,适合视频生成与数字人应用
由于真实世界中存在复杂且高度动态的运动,从多视角输入合成任意视角的动态视频极具挑战。基于神经辐射场或3D高斯泼溅的现有方法仅能建模细粒度运动,严重限制了应用范围。本文提出LocalDyGS,由两部分组成:1)将复杂动态场景分解为由种子定义的流线型局部空间,通过捕捉每个局部空间内的运动实现全局建模;2)对局部空间运动建模时解耦静态与动态特征——共享时间步的静态特征捕捉静态信息,动态残差场提供时序特异性特征,二者结合并解码生成时序高斯点,实现局部空间内运动建模。该方法构建了一种新颖的动态场景重建框架,不仅在多个细粒度数据集上表现优于当前最优(SOTA)方法,更首次实现了对更大、更复杂的高动态场景的建模。项目主页:https://wujh2001.github.io/LocalDyGS/
原文摘要 · Abstract (English)
Due to the complex and highly dynamic motions in the real world, synthesizing dynamic videos from multi-view inputs for arbitrary viewpoints is challenging. Previous works based on neural radiance field or 3D Gaussian splatting are limited to modeling fine-scale motion, greatly restricting their application. In this paper, we introduce LocalDyGS, which consists of two parts to adapt our method to both large-scale and fine-scale motion scenes: 1) We decompose a complex dynamic scene into streamlined local spaces defined by seeds, enabling global modeling by capturing motion within each local space. 2) We decouple static and dynamic features for local space motion modeling. A static feature shared across time steps captures static information, while a dynamic residual field provides time-specific features. These are combined and decoded to generate Temporal Gaussians, modeling motion within each local space. As a result, we propose a novel dynamic scene reconstruction framework to model highly dynamic real-world scenes more realistically. Our method not only demonstrates competitive performance on various fine-scale datasets compared to state-of-the-art (SOTA) methods, but also represents the first attempt to model larger and more complex highly dynamic scenes. Project page: https://wujh2001.github.io/LocalDyGS/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。