arXiv:2412.04447cs.AIcs.CV2024-12IJCV被引 37

评测多模态大模型在真实场景中的规划能力,发现其表现仍有巨大提升空间。

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

  • 构建基于第一视角视频的多场景规划评测基准,覆盖24个生活化任务
  • 21个主流多模态大模型在该基准上平均表现不佳,最高仅达60.3分
  • 提出无需训练的多模态思维链提示方法,使GPT-4V得分提升10.24分

多模态大语言模型(MLLM)虽在多模态理解与推理方面表现优异,但实现通用人工智能还需具备复杂环境下的有效规划能力。本文提出EgoPlan-Bench2,一个涵盖4大领域、24个真实生活场景的综合性评测基准,基于半自动流程结合第一人称视角视频构建,并经人工验证。该基准贴近人类日常问题解决方式。我们评估了21个先进MLLM,在真实场景规划任务中普遍表现有限,揭示出当前模型在复杂决策中的不足。为此,我们提出一种无需训练的多模态思维链(CoT)提示策略,通过分析多种多模态提示的有效性,在不增加训练成本的前提下,使GPT-4V在该基准上性能提升10.24分。本研究为未来提升多模态模型规划能力提供了关键洞见与工具。数据与代码已开源:https://qiulu66.github.io/egoplanbench2/

原文摘要 · Abstract (English)

The advent of Multimodal Large Language Models, leveraging the power of Large Language Models, has recently demonstrated superior multimodal understanding and reasoning abilities, heralding a new era for artificial general intelligence. However, achieving AGI necessitates more than just comprehension and reasoning. A crucial capability required is effective planning in diverse scenarios, which involves making reasonable decisions based on complex environments to solve real-world problems. Despite its importance, the planning abilities of current MLLMs in varied scenarios remain underexplored. In this paper, we introduce EgoPlan-Bench2, a rigorous and comprehensive benchmark designed to assess the planning capabilities of MLLMs across a wide range of real-world scenarios. EgoPlan-Bench2 encompasses everyday tasks spanning 4 major domains and 24 detailed scenarios, closely aligned with human daily life. EgoPlan-Bench2 is constructed through a semi-automatic process utilizing egocentric videos, complemented by manual verification. Grounded in a first-person perspective, it mirrors the way humans approach problem-solving in everyday life. We evaluate 21 competitive MLLMs and provide an in-depth analysis of their limitations, revealing that they face significant challenges in real-world planning. To further improve the planning proficiency of current MLLMs, we propose a training-free approach using multimodal Chain-of-Thought (CoT) prompting through investigating the effectiveness of various multimodal prompts in complex planning. Our approach enhances the performance of GPT-4V by 10.24 on EgoPlan-Bench2 without additional training. Our work not only sheds light on the current limitations of MLLMs in planning, but also provides insights for future enhancements in this critical area. We have made data and code available at https://qiulu66.github.io/egoplanbench2/.

多模态规划能力评测基准提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。