arXiv:2604.01678cs.CV2026-04

让动态场景的3D重建同时理解每个物体的身份和语义。

Director: Instance-aware Gaussian Splatting for Dynamic Scene Modeling and Understanding

  • 用可学习语义特征和语言模型监督高斯点,实现实例级建模。
  • 在真实动态场景中实现稳定追踪与开放词汇查询,精度优于现有方法。
  • 适合做视频理解、人形动画重建的研究者或开发者使用。

体素化视频旨在将动态场景建模为时间一致的4D表示。尽管基于高斯的方法在渲染质量上表现优异,但主要关注外观而忽略实例级结构,限制了高度动态场景中的稳定追踪与语义推理。本文提出Director,一种统一的时空高斯表示,联合建模人体动作、高保真渲染与实例级语义。核心思想是嵌入实例一致的语义能自然补充4D建模,实现更准确的场景分解并支持鲁棒动态理解。我们利用时序对齐的实例掩码与来自多模态大模型的句子嵌入,通过两个MLP解码器监督每个高斯点的可学习语义特征,实现语言对齐的4D表示并保证时间一致性。为增强时序稳定性,我们将2D光流与4D高斯点桥接并微调其运动,获得可靠初始化并减少漂移。训练中引入几何感知SDF约束及表面连续性正则项,提升动态前景建模的时间连贯性。实验表明,Director在实现时间一致的4D重建的同时,支持实例分割与开放词汇查询。

原文摘要 · Abstract (English)

Volumetric video seeks to model dynamic scenes as temporally coherent 4D representations. While recent Gaussian-based approaches achieve impressive rendering fidelity, they primarily emphasize appearance but are largely agnostic to instance-level structure, limiting stable tracking and semantic reasoning in highly dynamic scenarios. In this paper, we present Director, a unified spatio-temporal Gaussian representation that jointly models human performance, high-fidelity rendering, and instance-level semantics. Our key insight is that embedding instance-consistent semantics naturally complements 4D modeling, enabling more accurate scene decomposition while supporting robust dynamic scene understanding. To this end, we leverage temporally aligned instance masks and sentence embeddings derived from Multimodal Large Language Models to supervise the learnable semantic features of each Gaussian via two MLP decoders, enabling language-aligned 4D representations and enforcing identity consistency over time. To enhance temporal stability, we bridge 2D optical flow with 4D Gaussians and finetune their motions, yielding reliable initialization and reducing drift. For the training, we further introduce a geometry-aware SDF constraints, along with regularization terms that enforces surface continuity, enhancing temporal coherence in dynamic foreground modeling. Experiments demonstrate that Director achieves temporally coherent 4D reconstructions while simultaneously enabling instance segmentation and open-vocabulary querying.

动态建模高斯溅射实例分割语言对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。