构建多模态时间推理基准,评估视觉语言模型的规划能力
MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs
- 设计可扩展众包流程,构建带时序图标注的多模态食谱数据集
- 在6个主流视觉语言模型上验证其时序推理能力差异
- 适合研究多模态规划与复杂任务执行的学者使用
AI智能体需规划以实现涉及感知、子目标分解和执行的复杂目标。这些计划由按时间执行顺序(TEO)组织的有序步骤构成,形成有向无环图,确保每一步仅在前置条件满足后执行。现有研究对基础模型的时间执行理解局限于自动生成的标注、将TEO近似为线性链或仅文本输入。为填补此空白,我们提出MATEO(Multimodal Temporal Execution Order),一个用于评估和提升大型视觉语言模型(LVLMs)真实世界规划所需时间推理能力的基准。通过标准化编辑流程获取高质量专业多模态食谱语料库,将指令分解为离散步骤,并配以对应图像。设计并使用可扩展众包管道收集TEO图注释。利用MATEO,在六种前沿LVLM上评估了模型规模、语言上下文、多模态输入结构和微调策略的影响。
原文摘要 · Abstract (English)
AI agents need to plan to achieve complex goals that involve orchestrating perception, sub-goal decomposition, and execution. These plans consist of ordered steps structured according to a Temporal Execution Order (TEO, a directed acyclic graph that ensures each step executes only after its preconditions are satisfied. Existing research on foundational models' understanding of temporal execution is limited to automatically derived annotations, approximations of the TEO as a linear chain, or text-only inputs. To address this gap, we introduce MATEO (MultimodAl Temporal Execution Order), a benchmark designed to assess and improve the temporal reasoning abilities of Large Vision Language Models (LVLMs) required for real-world planning. We acquire a high-quality professional multimodal recipe corpus, authored through a standardized editorial process that decomposes instructions into discrete steps, each paired with corresponding images. We collect TEO annotations as graphs by designing and using a scalable crowdsourcing pipeline. Using MATEO, we evaluate six state-of-the-art LVLMs across model scales, varying language context, multimodal input structure, and fine-tuning strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。