arXiv:2505.12253cs.CV2025-05被引 20

让大模型理解4维场景中的动态物体,通过时空提示增强视觉表征。

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding

  • 用时空坐标生成动态感知的4D嵌入,替代固定空间提示。
  • 在4D场景理解任务中,动态物体识别准确率显著提升。
  • 适合需要理解动态物理世界的机器人、自动驾驶等应用。

尽管在2D图像理解上取得显著进展,大模态模型(LMMs)在真实世界中仍因缺乏空间表征能力而受限。现有3D LMM主要将3D位置作为固定空间提示嵌入视觉特征,但仅能处理静态背景,无法捕捉动态物体的时间变化。本文提出LLaVA-4D,一种通用的LMM框架,通过编码3D位置与1D时间生成动态感知的4D坐标嵌入,实现4D场景理解。我们证明,从视觉特征中解耦出的空间与时间成分更有效地区分背景与物体。因此,将4D时空提示嵌入这些特征可增强动态场景表征。通过对齐视觉时空嵌入与语言嵌入,LMM具备理解静态背景与动态物体的空间和时间特性的能力。此外,我们构建了一个带有时空坐标标注的4D视觉语言数据集,用于指令微调LMM。大量实验验证了该方法在不同4D场景理解任务中的有效性。

原文摘要 · Abstract (English)

Despite achieving significant progress in 2D image understanding, large multimodal models (LMMs) struggle in the physical world due to the lack of spatial representation. Typically, existing 3D LMMs mainly embed 3D positions as fixed spatial prompts within visual features to represent the scene. However, these methods are limited to understanding the static background and fail to capture temporally varying dynamic objects. In this paper, we propose LLaVA-4D, a general LMM framework with a novel spatiotemporal prompt for visual representation in 4D scene understanding. The spatiotemporal prompt is generated by encoding 3D position and 1D time into a dynamic-aware 4D coordinate embedding. Moreover, we demonstrate that spatial and temporal components disentangled from visual features are more effective in distinguishing the background from objects. This motivates embedding the 4D spatiotemporal prompt into these features to enhance the dynamic scene representation. By aligning visual spatiotemporal embeddings with language embeddings, LMMs gain the ability to understand both spatial and temporal characteristics of static background and dynamic objects in the physical world. Additionally, we construct a 4D vision-language dataset with spatiotemporal coordinate annotations for instruction fine-tuning LMMs. Extensive experiments have been conducted to demonstrate the effectiveness of our method across different tasks in 4D scene understanding.

4D理解时空建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。