arXiv:2503.10437cs.CV2025-03CVPR被引 40

用多模态大模型生成动态视频对象描述,实现4D场景下精准语言查询。

4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models

  • 通过多模态大模型生成逐帧物体级视频描述,替代传统视觉特征学习
  • 在多个基准上实现时间敏感与无关的开放词汇查询,精度与效率俱佳
  • 适合需要动态场景理解的智能交互、自动驾驶等应用

学习4D语言场以支持动态场景中时间敏感的开放式语言查询,对众多现实应用至关重要。尽管LangSplat成功将CLIP特征融入3D高斯表示,在静态3D场景中实现高精度与高效性,但无法处理动态4D场,因为CLIP专为静态图像-文本任务设计,难以捕捉视频中的时间动态。真实环境本质上是动态的,物体语义随时间演变。构建精确的4D语言场需获取像素对齐、物体级的视频特征,而当前视觉模型难以实现。为此,我们提出4D LangSplat,通过多模态大语言模型(MLLMs)直接从物体级视频描述生成的文本中学习4D语言场,实现对时间无关或时间敏感的开放词汇查询的高效处理。具体地,我们提出一种多模态物体级视频提示方法,结合视觉与文本提示,引导MLLM生成详尽、时间一致、高质量的视频物体描述。这些描述经大语言模型编码为高质量句子嵌入,作为像素对齐、物体特定的特征监督信号,通过共享嵌入空间支持开放词汇文本查询。考虑到4D场景中物体状态的平滑过渡,我们进一步提出状态可变形网络,有效建模随时间连续变化的过程。多基准实验结果表明,4D LangSplat在时间敏感与无关的开放词汇查询上均达到精确高效的性能。

原文摘要 · Abstract (English)

Learning 4D language fields to enable time-sensitive, open-ended language queries in dynamic scenes is essential for many real-world applications. While LangSplat successfully grounds CLIP features into 3D Gaussian representations, achieving precision and efficiency in 3D static scenes, it lacks the ability to handle dynamic 4D fields as CLIP, designed for static image-text tasks, cannot capture temporal dynamics in videos. Real-world environments are inherently dynamic, with object semantics evolving over time. Building a precise 4D language field necessitates obtaining pixel-aligned, object-wise video features, which current vision models struggle to achieve. To address these challenges, we propose 4D LangSplat, which learns 4D language fields to handle time-agnostic or time-sensitive open-vocabulary queries in dynamic scenes efficiently. 4D LangSplat bypasses learning the language field from vision features and instead learns directly from text generated from object-wise video captions via Multimodal Large Language Models (MLLMs). Specifically, we propose a multimodal object-wise video prompting method, consisting of visual and text prompts that guide MLLMs to generate detailed, temporally consistent, high-quality captions for objects throughout a video. These captions are encoded using a Large Language Model into high-quality sentence embeddings, which then serve as pixel-aligned, object-specific feature supervision, facilitating open-vocabulary text queries through shared embedding spaces. Recognizing that objects in 4D scenes exhibit smooth transitions across states, we further propose a status deformable network to model these continuous changes over time effectively. Our results across multiple benchmarks demonstrate that 4D LangSplat attains precise and efficient results for both time-sensitive and time-agnostic open-vocabulary queries.

4D重建语言模型动态场景多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。