用事件相机与高动态场景分离建模,提升动态物体重建精度。
STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
- 通过帧-事件融合分离背景与动态物体的时空特征。
- 在动态场景中实现时间连续的物体运动重建,有效缓解时空错配。
- 适合需要高精度动态场景重建的研究者与工业应用。
高动态场景重建旨在用刚性空间特征表示静态背景,用连续变形的时空特征表示动态物体。现有方法通常采用统一的高斯表示模型(如高斯)直接从帧图像中匹配动态场景的时空特征,但该统一范式因帧成像导致的潜在时间不连续性,以及背景与物体间空间特征异质性而失效。为此,本文将事件相机与帧相机结合,提出一种时空解耦的高斯点阵框架以实现高动态场景重建。研究发现,基于帧的特征中背景与物体存在外观差异,基于事件的特征中二者存在运动差异,这促使我们通过聚类区分其时空特征。进一步发现,高斯表示与事件数据共享一致的时空特性,可作为先验指导物体高斯的时空解耦。在高斯点阵框架内,逐步实现场景-物体解耦,提升背景与物体间的时空判别能力,从而渲染出时间连续的动态场景。大量实验验证了所提方法的优越性。
原文摘要 · Abstract (English)
High-dynamic scene reconstruction aims to represent static background with rigid spatial features and dynamic objects with deformed continuous spatiotemporal features. Typically, existing methods adopt unified representation model (e.g., Gaussian) to directly match the spatiotemporal features of dynamic scene from frame camera. However, this unified paradigm fails in the potential discontinuous temporal features of objects due to frame imaging and the heterogeneous spatial features between background and objects. To address this issue, we disentangle the spatiotemporal features into various latent representations to alleviate the spatiotemporal mismatching between background and objects. In this work, we introduce event camera to compensate for frame camera, and propose a spatiotemporal-disentangled Gaussian splatting framework for high-dynamic scene reconstruction. As for dynamic scene, we figure out that background and objects have appearance discrepancy in frame-based spatial features and motion discrepancy in event-based temporal features, which motivates us to distinguish the spatiotemporal features between background and objects via clustering. As for dynamic object, we discover that Gaussian representations and event data share the consistent spatiotemporal characteristic, which could serve as a prior to guide the spatiotemporal disentanglement of object Gaussians. Within Gaussian splatting framework, the cumulative scene-object disentanglement can improve the spatiotemporal discrimination between background and objects to render the time-continuous dynamic scene. Extensive experiments have been performed to verify the superiority of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。