构建桥梁巡检问答基准,推动具身智能理解长程空间信息
BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
- 以真实桥梁巡检报告为依据,构建200个场景的开放词汇问答数据集
- 提出图像引用相关性新指标,发现现有模型存在显著性能差距
- 设计基于马尔可夫决策过程的视觉推理框架,提升多尺度空间理解能力
在真实世界中部署能回答环境问题的具身智能体仍面临挑战,部分原因在于缺乏用于情景记忆具身问答(EQA)的基准。受基础设施巡检难题启发,我们提出巡检EQA作为推进情景记忆EQA的关键问题类别。该任务需要多尺度推理和长程空间理解,同时具备标准化评估、专业巡检报告作为语义锚点以及第一人称图像输入。我们引入BridgeEQA,一个包含2,200个开放词汇问答对(类似OpenEQA风格)的基准,基于200个真实桥梁场景的巡检报告,每场景平均47.93张图像。我们进一步提出新指标Image Citation Relevance,用于评估模型引用相关图像的能力。对主流视觉语言模型的评估显示显著性能差距。为此,我们提出Embodied Memory Visual Reasoning(EMVR),将巡检EQA任务建模为马尔可夫决策过程,实验表明其优于基线模型。代码与数据集已公开于https://drags99.github.io/bridge-eqa/
原文摘要 · Abstract (English)
Deploying embodied agents that can answer questions about their surroundings in realistic real-world settings remains difficult, partly due to the scarcity of benchmarks for episodic memory Embodied Question Answering (EQA). Inspired by the challenges of infrastructure inspections, we propose Inspection EQA as a compelling problem class for advancing episodic memory EQA. It demands multi-scale reasoning and long-range spatial understanding, while offering standardized evaluation, professional inspection reports as grounding, and egocentric imagery. We introduce BridgeEQA, a benchmark of 2,200 open-vocabulary question-answer pairs (in the style of OpenEQA) grounded in professional inspection reports across 200 real-world bridge scenes with 47.93 images on average per scene. We further propose a new EQA metric Image Citation Relevance to evaluate the ability of a model to cite relevant images. Evaluations of state-of-the-art vision-language models reveal substantial performance gaps. To address this, we propose Embodied Memory Visual Reasoning (EMVR), which formulates the inspection EQA task as a Markov decision process. EMVR shows strong performance over the baselines. Code and dataset are available at https://drags99.github.io/bridge-eqa/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。