arXiv:2509.23828cs.CV2025-09被引 10

首个统一4D场景理解与生成的视觉语言模型,支持时空感知的联合推理。

Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation

  • 通过自适应交叉注意力融合语义、外观与4D几何特征,构建统一的时空表示
  • 在多个基准上达到或超过现有最优模型性能,实现理解与生成的协同优化
  • 适合需要动态场景建模的研究者,尤其关注4D内容生成与多任务理解

视觉语言模型(VLM)在2D场景理解与生成方面表现优异,但将其统一扩展至真实物理世界仍面临挑战。现有3D和4D方法通常将场景几何嵌入自回归模型用于语义理解,扩散模型用于内容生成,这种范式差异导致单一模型难以同时处理两类任务,尤其在动态4D场景中,时空建模至关重要。本文提出Uni4D-LLM,首个具备时空感知能力的统一视觉语言框架,用于4D场景的理解与生成。设计基于两大核心洞察:1)统一需共享表示——提取语义特征用于理解,噪声注入的外观特征用于生成,结合4D几何线索,通过自适应交叉注意力融合为时空感知的视觉表示;2)统一需共享架构——自回归与扩散均基于Transformer骨干网络,可集成于单一大语言模型,配合任务特定头实现端到端联合预测。通过在多样化的4D视觉语言数据集上进行指令微调,提升跨任务泛化能力。大量实验表明,Uni4D-LLM在多个基准上达到或优于当前最优模型,首次实现了4D场景理解与生成的真正统一。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have demonstrated strong performance in 2D scene understanding and generation, but extending this unification to the physical world remains an open challenge. Existing 3D and 4D approaches typically embed scene geometry into autoregressive model for semantic understanding and diffusion model for content generation. This paradigm gap prevents a single model from jointly handling both tasks, especially in dynamic 4D settings where spatiotemporal modeling is critical. We propose Uni4D-LLM, the first unified VLM framework with spatiotemporal awareness for 4D scene understanding and generation. Our design is guided by two key insights: 1) Unification requires a shared representation. We extract semantic features for understanding and noisy-injected appearance features for generation, incorporate 4D geometric cues, and fuse them into a spatiotemporal-aware visual representation through adaptive cross-attention. 2) Unification requires a shared architecture. Both autoregression and diffusion are built on Transformer backbones, and this enables integration into a single LLM with task-specific heads. By aligning visual and linguistic representations, our Uni4D-LLM produces predictions for both understanding and generation within one Transformer-based framework. We further apply instruction fine-tuning on diverse 4D vision-language datasets to improve generalization across tasks. Extensive experiments on multiple benchmarks demonstrate that Uni4D-LLM achieves competitive or superior results compared to state-of-the-art models and offers the first true unification of 4D scene understanding and generation.

4D生成视觉语言模型时空建模统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。