arXiv:2608.18734cs.CV2026-08中稿 · ECCV

首个直接处理动态点云的4D视觉语言模型,实现时空几何与语言对齐。

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

论文配图:CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
图 1 · 摘自论文原文
  • 基于对比学习,将动态点云的时空几何特征与自然语言对齐
  • 在4D动作基准上性能提升约16.75%,超越现有方法
  • 适用于需要理解动态场景的机器人、自动驾驶等应用

4D理解与推理是嵌入式智能体在动态物理环境中运行的基础能力。然而,现有视觉编码器大多局限于静态2D图像或3D点云,缺乏时间建模,或仅基于2D视频但缺乏精确的几何深度推理。因此,当前方法无法联合捕捉动态场景中的空间结构与运动演化。本文提出CL4D,首个直接作用于动态点云的基础4D视觉编码器,通过对比学习目标将时空几何表征与自然语言描述对齐。通过学习文本与4D场景动态之间的共享嵌入空间,CL4D实现了动态环境下的零样本运动到文本及文本到运动检索,并可作为下游4D视觉语言任务的基础编码器。在此基础上,我们构建了4DVLM,一个直接基于动态几何表示生成语言的4D视觉语言模型。4DVLM是首个不依赖2D图像、2D视频或静态3D点云的视觉语言模型。我们在新构建的DynAction4D数据集上训练CL4D和4DVLM,该数据集涵盖多样人体动作及其与不同物体交互和场景环境的变化。在多个4D人体动作基准上的实验表明,CL4D性能达到当前最优,较之前方法提升约16.75%。此外,即使为前沿视频视觉语言模型(如Gemini、GPT-5)提供与4DVLM相同场景的RGB视频序列,4DVLM仍表现更优。

原文摘要 · Abstract (English)

4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited to static 2D images or 3D point clouds without temporal modeling, or to 2D videos that lack accurate geometric depth reasoning. Consequently, current approaches fail to jointly capture spatial structure and motion evolution in dynamic scenes. We present CL4D, the first foundational 4D vision encoder that directly operates on dynamic point clouds, trained with a contrastive learning objective to align spatio-temporal geometric representations with natural language descriptions. By learning a shared embedding space between text and 4D scene dynamics, CL4D enables zero-shot motion-to-text and text-to-motion retrieval in dynamic environments and serves as a foundational 4D vision encoder for downstream 4D vision-language tasks. Building on this encoder, we introduce 4DVLM, a 4D vision-language model that conditions language generation on dynamic geometric representations. 4DVLM is the first VLM designed to operate directly on 4D point clouds without relying on 2D images, 2D videos, or static 3D point clouds. We train CL4D and subsequently 4DVLM on a newly constructed dataset termed DynAction4D capturing diverse human motions across varying object interactions and scene environments. Extensive experiments across multiple 4D human action benchmarks demonstrate that CL4D achieves state-of-the-art performance, with improvements of approximately ~16.75% over prior methods. Furthermore, 4DVLM outperforms frontier video VLMs such as Gemini and GPT-5 even when these models are provided with RGB video sequences corresponding to the same scenes represented as 4D point clouds for 4DVLM.

4D视觉动态点云视觉语言对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。