arXiv:2509.19002cs.CVcs.AI2025-09AAAI

用旅行视频重建任务评估大模型时空理解能力

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

  • 设计200段旅行视频,通过行程重构考验模型时空推理
  • 顶尖模型在长距离跨区域视频中表现不佳,准确率不高
  • 可指导智能旅行助手开发,适合导航与具身智能研究者

多模态大语言模型(MLLMs)在视频理解方面取得显著进展,但现有基准大多聚焦室内场景或短时户外活动,对长距离旅行所涉及的时空挑战关注不足。掌握长程地理-时间轨迹对下一代MLLMs至关重要,支撑具身AI规划与导航等实际应用。为此,我们提出VIR-Bench,一个包含200段旅行视频的新基准,将行程重建作为核心任务,用于评估和推动MLLMs的时空智能发展。实验表明,包括闭源在内的先进模型在该任务上得分普遍偏低,凸显其在长时空尺度视频处理上的困难。进一步地,我们基于VIR-Bench分析结果构建原型旅行规划代理,其推荐质量显著提升,验证了该评测框架不仅能有效衡量模型性能,还能转化为实际应用中的性能增益。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet current video benchmarks focus largely on indoor scenes or short-range outdoor activities, leaving the challenges associated with long-distance travel largely unexplored. Mastering extended geospatial-temporal trajectories is critical for next-generation MLLMs, underpinning real-world tasks such as embodied-AI planning and navigation. To bridge this gap, we present VIR-Bench, a novel benchmark consisting of 200 travel videos that frames itinerary reconstruction as a challenging task designed to evaluate and push forward MLLMs' geospatial-temporal intelligence. Experimental results reveal that state-of-the-art MLLMs, including proprietary ones, struggle to achieve high scores, underscoring the difficulty of handling videos that span extended spatial and temporal scales. Moreover, we conduct an in-depth case study in which we develop a prototype travel-planning agent that leverages the insights gained from VIR-Bench. The agent's markedly improved itinerary recommendations verify that our evaluation protocol not only benchmarks models effectively but also translates into concrete performance gains in user-facing applications.

多模态模型时空理解旅行规划视频评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。