构建细粒度诊断基准,评估智能体在真实场景中的推理能力。
ERQA-Plus: A Diagnostic Benchmark for Reasoning in Embodied AI

- 按感知、动作、社交等五大类设计推理任务,确保可解释性。
- 最强模型仅达83.4%准确率,空间与意图推理仍存明显短板。
- 适合研究通用智能体推理机制的学者和开发者使用。
通用具身智能体不仅需识别物体,还需基于情境视觉观察进行空间关系、行为过程、人类意图、环境约束及常识后果的推理。现有视觉与具身问答数据集对推理依赖的控制有限,难以区分真实推理与捷径匹配。本文提出ERQA-Plus,一个具身推理诊断基准,包含1,766个问题-答案实例,基于711张机器人视角图像,按感知、动作、社会互动、导航-环境与上下文常识五大类组织。数据通过多阶段生成与验证流程构建:基于分类体系生成问题,自动质量判断,迭代修订,人工评估,提升视觉锚定性、答案有效性与推理质量。我们评测了包括LLaVA-NeXT-8B、Prismatic-7B、MiniCPM-V-4.5-8B、Qwen3-VL、RoboRefer-8B、RoboBrain2.5-8B在内的代表性模型。尽管最强模型Qwen3-VL-32B达到83.4%总体准确率与61.4分SBERT得分,但在空间推理、程序推理、事件预测与意图推断上仍存在持续弱项。该基准为衡量智能体在哪些推理形式上可靠提供细粒度评估框架。数据集已公开于https://huggingface.co/datasets/huggingdas/erqa-plus,项目页见https://github.com/LUNAProject22/erqa-plus。
原文摘要 · Abstract (English)
Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and commonsense consequences from situated visual observations. Yet existing visual and embodied question answering benchmarks often provide limited control over the reasoning dependencies being tested, making it difficult to distinguish grounded embodied reasoning from shortcut-driven visual or linguistic pattern matching. We present ERQA-Plus, a diagnostic benchmark for reasoning in embodied AI. ERQA-Plus contains 1,766 question-answer instances grounded in 711 robot-centric images and organized according to a structured taxonomy spanning perceptual, action-centric, social-interaction, navigation-environmental, and contextual commonsense reasoning. The dataset is constructed using a multi-stage generation and validation pipeline that combines taxonomy-guided question generation, automatic quality judging, iterative revision, and human assessment to improve visual grounding, answer validity, and reasoning quality. We benchmark representative general-purpose vision-language models and embodied models, including LLaVA-NeXT-8B, Prismatic-7B, MiniCPM-V-4.5-8B, Qwen3-VL, RoboRefer-8B, and RoboBrain2.5-8B. Although the strongest model, Qwen3-VL-32B, achieves 83.4% overall accuracy and 61.4 SBERT score, category-level results reveal persistent weaknesses in spatial reasoning, procedural reasoning, event prediction, and intention inference. ERQA-Plus therefore provides a fine-grained evaluation framework for measuring not only whether embodied agents answer correctly, but also which forms of embodied reasoning they can and cannot perform reliably. The dataset is available https://huggingface.co/datasets/huggingdas/erqa-plus and the project page at https://github.com/LUNAProject22/erqa-plus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。