arXiv:2605.25813cs.RO2026-05

构建首个覆盖四维度的具身问答数据集,推动智能体决策研究

Extending Embodied Question Answering from Perception to Decision

论文配图:Extending Embodied Question Answering from Perception to Decision
图 1 · 摘自论文原文
  • 设计四维具身推理框架,涵盖场景、空间、任务动态与即时决策
  • 包含超400万问答对,支持多层次标注与跨场景评估
  • 提出统一基准模型RoboDecision,整合感知、推理与行动决策

具身问答(EQA)将感知、推理与交互整合于具身环境中。然而,现有数据集与评测标准仍碎片化,仅聚焦空间理解或流程推理等有限能力,缺乏统一的大规模评估框架。本文提出EQA-Decision,一个大规模具身问答数据集,系统覆盖四个互补的具身推理维度:静态场景构建、空间理解、任务动态推理与即时决策。数据集包含超过四百万个问题-答案对,并在多样化的具身场景中提供分层标注。此外,我们开发了RoboDecision模型,作为与EQA-Decision基准对齐的强基线,提供统一框架,联合评估具身环境中的感知、推理与动作级决策。结果表明,EQA-Decision有效评估并提升视觉语言模型在空间与交互推理方面的能力,为推进具身智能研究奠定坚实基础。

原文摘要 · Abstract (English)

Embodied Question Answering (EQA) connects perception, reasoning, and interaction within embodied environments. However, existing datasets and benchmarks remain fragmented, each focusing on a limited subset of reasoning skills such as spatial understanding or procedural reasoning, without offering a unified large-scale framework for comprehensive evaluation. We present EQA-Decision, a large-scale embodied QA dataset that systematically covers four complementary dimensions of embodied reasoning: static scene construction, spatial understanding, task dynamics reasoning, and instant decision. The dataset contains over four million question-answer pairs with hierarchical annotations across diverse embodied scenarios. In addition, we develop RoboDecision, a strong baseline model aligned with the EQA-Decision Benchmark, providing a unified framework that jointly evaluates perception, reasoning, and action-level decision-making in embodied environments. Results demonstrate that EQA-Decision effectively benchmarks and enhances VLM capabilities in spatial and interaction reasoning, providing a solid foundation for advancing embodied intelligence research.

具身智能问答系统决策推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。