arXiv:2503.06157cs.CVcs.AI2025-03ACL被引 49

评测视频大模型在城市环境中的沉浸式认知能力,发现其在推理与导航上仍有明显不足。

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

  • 构建真实城市与模拟环境的1.5千段第一人称视频数据集。
  • 生成5.2千道多选题,评估模型在回忆、感知、推理与导航上的表现。
  • 揭示因果推理与导航能力强相关,适合研究具身智能的学者参考。

大型多模态模型展现出强大智能,但其在开放城市三维空间中移动时的具身认知能力仍待探索。我们提出一个基准,用于评估视频-大语言模型(Video-LLMs)是否能像人类一样自然处理连续的第一人称视觉观测,实现记忆、感知、推理与导航。通过手动操控无人机采集真实城市与模拟环境的3D具身运动视频数据,共获得1.5千段视频片段,并设计流程生成5.2千道多选题。对17种主流Video-LLMs的评估揭示了当前在城市具身认知方面的局限性。相关性分析显示,因果推理与记忆、感知、导航能力高度相关,而反事实与关联推理与其他任务相关性较低。同时,通过微调验证了城市具身任务中从仿真到真实的迁移潜力。

原文摘要 · Abstract (English)

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored. We introduce a benchmark to evaluate whether video-large language models (Video-LLMs) can naturally process continuous first-person visual observations like humans, enabling recall, perception, reasoning, and navigation. We have manually control drones to collect 3D embodied motion video data from real-world cities and simulated environments, resulting in 1.5k video clips. Then we design a pipeline to generate 5.2k multiple-choice questions. Evaluations of 17 widely-used Video-LLMs reveal current limitations in urban embodied cognition. Correlation analysis provides insight into the relationships between different tasks, showing that causal reasoning has a strong correlation with recall, perception, and navigation, while the abilities for counterfactual and associative reasoning exhibit lower correlation with other tasks. We also validate the potential for Sim-to-Real transfer in urban embodiment through fine-tuning.

具身智能视频理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。