arXiv:2503.11089cs.ROcs.AI2025-03被引 9

用动态场景图引导思维链,让机器人更好理解复杂空间任务

EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

  • 构建动态场景图,结构化表示空间关系
  • 零样本推理准确率显著优于现有方法
  • 适合需要长期交互的智能体空间决策场景

尽管多模态大语言模型在具身智能领域取得突破,但在复杂长时程任务中的空间推理仍面临挑战。为此,我们提出EmbodiedVSR框架,通过动态场景图引导思维链(CoT)推理,增强具身代理的空间理解能力。该方法通过动态构建结构化知识表征,实现无需任务微调的零样本空间推理,不仅解耦复杂的空间关系,还使推理步骤与环境动态行为对齐。为严格评估性能,我们引入eSpatial-Benchmark数据集,包含真实具身场景、细粒度空间标注及可变难度任务。实验表明,该框架在长时程任务中显著提升准确率与推理连贯性,揭示了具备结构化可解释推理机制的多模态大模型在具身智能中的巨大潜力,为实际空间应用提供可靠路径。代码与数据集即将开源。

原文摘要 · Abstract (English)

While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex long-horizon tasks. To address this gap, we propose EmbodiedVSR (Embodied Visual Spatial Reasoning), a novel framework that integrates dynamic scene graph-guided Chain-of-Thought (CoT) reasoning to enhance spatial understanding for embodied agents. By explicitly constructing structured knowledge representations through dynamic scene graphs, our method enables zero-shot spatial reasoning without task-specific fine-tuning. This approach not only disentangles intricate spatial relationships but also aligns reasoning steps with actionable environmental dynamics. To rigorously evaluate performance, we introduce the eSpatial-Benchmark, a comprehensive dataset including real-world embodied scenarios with fine-grained spatial annotations and adaptive task difficulty levels. Experiments demonstrate that our framework significantly outperforms existing MLLM-based methods in accuracy and reasoning coherence, particularly in long-horizon tasks requiring iterative environment interaction. The results reveal the untapped potential of MLLMs for embodied intelligence when equipped with structured, explainable reasoning mechanisms, paving the way for more reliable deployment in real-world spatial applications. The codes and datasets will be released soon.

具身智能空间推理思维链场景图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。