构建可分级难度的机器人视觉推理基准,支持动态评估模型能力。
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- 定义推理复杂度并设计自适应提问引擎生成多级问题。
- 扩展JRDB数据集,加入人-物交互与几何关系标注。
- 适合研究机器人视觉推理与多步骤认知评估的团队使用。
视觉语言模型(VLM)和大语言模型(LLM)的进步显著提升了机器人等具身智能体的视觉推理能力。然而,现有视觉推理基准普遍存在推理复杂度定义模糊、无法控制生成不同难度的问题、缺乏任务定制化以及缺少结构化分步推理标注(工作流)等问题。为此,我们正式定义了推理复杂度,提出自适应查询引擎,可生成具有不同复杂度的可定制问题,并附带详细中间步骤标注;同时在JRDB数据集基础上,增加人-物交互与几何关系标注,构建了适用于人密集环境下的视觉推理基准JRDB-Reasoning。该引擎与基准支持对视觉推理框架进行细粒度评估,并实现对视觉-语言模型在不同推理层级上的动态测评。
原文摘要 · Abstract (English)
Recent advances in Vision-Language Models (VLMs) and large language models (LLMs) have greatly enhanced visual reasoning, a key capability for embodied AI agents like robots. However, existing visual reasoning benchmarks often suffer from several limitations: they lack a clear definition of reasoning complexity, offer have no control to generate questions over varying difficulty and task customization, and fail to provide structured, step-by-step reasoning annotations (workflows). To bridge these gaps, we formalize reasoning complexity, introduce an adaptive query engine that generates customizable questions of varying complexity with detailed intermediate annotations, and extend the JRDB dataset with human-object interaction and geometric relationship annotations to create JRDB-Reasoning, a benchmark tailored for visual reasoning in human-crowded environments. Our engine and benchmark enable fine-grained evaluation of visual reasoning frameworks and dynamic assessment of visual-language models across reasoning levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。