arXiv:2509.21662cs.LG2025-09EMNLP

零样本生成图文一致的步骤规划,靠物体状态推理链提升准确性

MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning

  • 用物体状态推理链显式建模物品变化过程
  • 图文对齐率提升11.9%,步骤顺序准确率提高26.7%
  • 适合需要跨模态一致性规划的应用场景

多模态程序规划(MPP)旨在生成融合文本与图像的分步指令,核心挑战在于保持跨模态的物体状态一致性并生成有信息量的计划。现有方法常使用大语言模型(LLMs)优化文本步骤,但视觉物体状态对齐与系统性评估仍研究不足。我们提出MMPlanner,一种零样本多模态程序规划框架,引入物体状态推理思维链(OSR-CoT)提示,显式建模物体状态转移,生成精准的多模态计划。为评估计划质量,设计基于大语言模型的评判协议,用于衡量规划准确性和跨模态对齐;进一步提出视觉步骤重排任务以测量时间连贯性。在RECIPEPLAN和WIKIPLAN数据集上的实验表明,MMPlanner达到当前最优性能,文本规划准确率提升6.8%,跨模态对齐提升11.9%,视觉步骤排序准确率提升26.7%。

原文摘要 · Abstract (English)

Multimodal Procedural Planning (MPP) aims to generate step-by-step instructions that combine text and images, with the central challenge of preserving object-state consistency across modalities while producing informative plans. Existing approaches often leverage large language models (LLMs) to refine textual steps; however, visual object-state alignment and systematic evaluation are largely underexplored. We present MMPlanner, a zero-shot MPP framework that introduces Object State Reasoning Chain-of-Thought (OSR-CoT) prompting to explicitly model object-state transitions and generate accurate multimodal plans. To assess plan quality, we design LLM-as-a-judge protocols for planning accuracy and cross-modal alignment, and further propose a visual step-reordering task to measure temporal coherence. Experiments on RECIPEPLAN and WIKIPLAN show that MMPlanner achieves state-of-the-art performance, improving textual planning by +6.8%, cross-modal alignment by +11.9%, and visual step ordering by +26.7%

多模态规划思维链物体状态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。