arXiv:2410.03907cs.CL2024-10EMNLP被引 14

构建1187个家庭活动的多模态规划基准,评估视觉语言模型的推理能力

ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities

  • 基于ChatGPT与iGibson2模拟器构建多模态任务数据集
  • 当前视觉语言模型在正常与反事实场景中仍无法生成人类级计划
  • 提供自动评估指标,支持未来对规划能力的研究

大型语言模型(LLMs)因强大的推理能力被用于处理文本任务描述并完成具身智能中的程序性规划。然而,关于视觉语言模型(VLMs)在多模态输入下的表现仍缺乏研究。反事实规划(评估模型对替代任务情境的推理能力)也未被充分探索。为评估多模态和反事实两个维度的规划能力,我们提出ActPlan-1K。该基准基于ChatGPT与家庭活动仿真器iGibson2构建,包含153项活动、1,187个实例。每个实例包含自然语言任务描述及模拟器生成的多张环境图像,真实计划为场景中物体的动作序列。在典型VLM上同时评估正确性和常识合理性。结果表明,当前VLM在正常及反事实活动的规划上仍难以达到人类水平。我们进一步通过微调BLEURT模型,提供自动评估指标以促进后续研究。

原文摘要 · Abstract (English)

Large language models~(LLMs) have been adopted to process textual task description and accomplish procedural planning in embodied AI tasks because of their powerful reasoning ability. However, there is still lack of study on how vision language models~(VLMs) behave when multi-modal task inputs are considered. Counterfactual planning that evaluates the model's reasoning ability over alternative task situations are also under exploited. In order to evaluate the planning ability of both multi-modal and counterfactual aspects, we propose ActPlan-1K. ActPlan-1K is a multi-modal planning benchmark constructed based on ChatGPT and household activity simulator iGibson2. The benchmark consists of 153 activities and 1,187 instances. Each instance describing one activity has a natural language task description and multiple environment images from the simulator. The gold plan of each instance is action sequences over the objects in provided scenes. Both the correctness and commonsense satisfaction are evaluated on typical VLMs. It turns out that current VLMs are still struggling at generating human-level procedural plans for both normal activities and counterfactual activities. We further provide automatic evaluation metrics by finetuning over BLEURT model to facilitate future research on our benchmark.

视觉语言模型程序规划具身智能多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。