arXiv:2507.23382cs.CLcs.AI2025-07中稿 · ACM Multimedia 202…被引 3

首个评估多模态大模型复杂约束规划能力的基准测试

MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models

  • 构建真实场景下的三类多模态规划任务,引入分级约束
  • 闭源模型仅21.3%可行计划,开源模型平均低于11%
  • 揭示模型对约束敏感性,适合研究多模态推理与规划

多模态规划能力指在多模态上下文中预测、推理并设计多步任务执行流程的能力,对复杂推理与决策至关重要。然而现有基准存在两大问题:(1) 无法直接评估真实世界多模态规划能力;(2) 缺乏跨模态的显式或隐式约束。为此,我们提出多模态复杂约束规划基准(MPCC),首次系统评估多模态大语言模型(MLLMs)处理多模态约束的规划能力。为解决第一问题,MPCC聚焦飞行规划、日程规划和会议规划三类真实任务;为解决第二问题,引入预算、时间、空间等复杂约束,并设三级难度(EASY、MEDIUM、HARD)以分离约束复杂度与搜索空间扩张。13个先进MLLMs实验表明:闭源模型仅21.3%生成可行计划,开源模型平均低于11%。此外,发现MLLMs对约束复杂度高度敏感,传统多模态提示策略在多约束场景中失效。本工作形式化了多模态规划中的约束,提供严谨评估框架,凸显真实应用中需发展约束感知推理能力。

原文摘要 · Abstract (English)

Multimodal planning capabilities refer to the ability to predict, reason, and design steps for task execution with multimodal context, which is essential for complex reasoning and decision-making across multiple steps. However, current benchmarks face two key challenges: (1) they cannot directly assess multimodal real-world planning capabilities, and (2) they lack constraints or implicit constraints across modalities. To address these issues, we introduce Multimodal Planning with Complex Constraints (MPCC), the first benchmark to systematically evaluate MLLMs' ability to handle multimodal constraints in planning. To address the first challenge, MPCC focuses on three real-world tasks: Flight Planning, Calendar Planning, and Meeting Planning. To solve the second challenge, we introduce complex constraints (e.g. budget, temporal, and spatial) in these tasks, with graded difficulty levels (EASY, MEDIUM, HARD) to separate constraint complexity from search space expansion. Experiments on 13 advanced MLLMs reveal significant challenges: closed-source models achieve only 21.3% feasible plans, while open-source models average below 11%. Additionally, we observe that MLLMs are highly sensitive to constraint complexity and that traditional multimodal prompting strategies fail in multi-constraint scenarios. Our work formalizes multimodal constraints in planning, provides a rigorous evaluation framework, and highlights the need for advancements in constraint-aware reasoning for real-world MLLM applications.

多模态规划约束推理基准测试MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。