arXiv:2605.04527cs.CV2026-05

用动态点云学习4D物体的几何与外观表示,高效且只需简单输入。

Velox: Learning Representations of 4D Geometry and Appearance

论文配图:Velox: Learning Representations of 4D Geometry and Appearance
图 1 · 摘自论文原文
  • 通过编码器将时空点云压缩为可动态变化的形状令牌。
  • 在视频到4D生成等三项任务中表现优异,验证了表示的有效性。
  • 仅需无结构动态点云即可构建,适合实时或低资源场景使用。

我们提出一种学习4D物体隐式表示的框架,该表示具备描述性(忠实捕捉几何与外观)、压缩性(提升下游效率)和可访问性(仅需无结构动态点云作为输入)。Velox训练一个编码器,将时空颜色点云压缩为一组动态形状令牌。这些令牌通过两个互补解码器进行监督:4D表面解码器建模随时间变化的表面分布以捕捉几何;高斯解码器将令牌映射为3D高斯,辅助学习外观。为验证表示的实用性,我们在三个下游任务上评估:视频到4D生成、3D跟踪和图像到4D生成的布料模拟,结果在所有场景中均表现良好。

原文摘要 · Abstract (English)

We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal input, i.e., an unstructured dynamic point cloud, to construct. Specifically, Velox trains an encoder to compress spatiotemporal color point clouds into a set of dynamic shape tokens. These tokens are supervised using two complementary decoders: a 4D surface decoder, which models the time-varying surface distribution capturing the geometry; and a Gaussian decoder, which maps the tokens to 3D Gaussians, helping learn appearance. To demonstrate the utility of our representation, we evaluate it across three downstream tasks -- video-to-4D generation, 3D tracking, and cloth simulation via image-to-4D generation -- and observe strong performances in all settings.

4D重建点云表示动态形状

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。