arXiv:2410.10719cs.CVcs.GR2024-10被引 14

用语言控制动态3D场景,实现视频中事件的时空定位

4-LEGS: 4D Language Embedded Gaussian Splatting

  • 将3D高斯点云扩展为4D时序表示,融合语言嵌入
  • 支持从文本指令中精准定位视频里的动作发生位置与时间
  • 适用于人和动物行为视频分析,交互式语义操控

神经表示的兴起彻底改变了我们数字化呈现3D场景的方式,实现了从新视角合成逼真图像。近年来,已有多种方法将低层表示与场景中的高层语义理解相连接,把2D图像中的丰富语义迁移到3D空间,将高维空间特征压缩到3D结构中。本文关注如何将语言与世界的动态建模相结合。我们展示了如何基于3D高斯点阵(3D Gaussian Splatting)构建4D时空表示,从而实现一个交互式界面,用户可通过文本提示在视频中精确定位事件的时空位置。我们在公开的3D视频数据集上验证了系统性能,涵盖人类和动物执行各种动作的场景。

原文摘要 · Abstract (English)

The emergence of neural representations has revolutionized our means for digitally viewing a wide range of 3D scenes, enabling the synthesis of photorealistic images rendered from novel views. Recently, several techniques have been proposed for connecting these low-level representations with the high-level semantics understanding embodied within the scene. These methods elevate the rich semantic understanding from 2D imagery to 3D representations, distilling high-dimensional spatial features onto 3D space. In our work, we are interested in connecting language with a dynamic modeling of the world. We show how to lift spatio-temporal features to a 4D representation based on 3D Gaussian Splatting. This enables an interactive interface where the user can spatiotemporally localize events in the video from text prompts. We demonstrate our system on public 3D video datasets of people and animals performing various actions.

4D建模语言控制高斯点云视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。