评测模型在视觉反馈下调整高阶行动计划的能力
AsgardBench -- Evaluating Visually Grounded Interactive Planning Under Minimal Feedback
- 仅提供图像、动作历史和成败信号,模拟真实交互中的计划修正
- 108个任务实例覆盖12类场景,通过环境变化诱导不同执行路径
- 揭示大模型缺乏视觉感知与状态追踪能力,导致计划失效
AsgardBench旨在评估视觉引导的高层次动作序列生成与交互式规划能力,重点关注执行过程中基于视觉观察对计划的动态调整,而非导航或底层操作。该基准聚焦于交互式规划能力,既高于离线高层规划,又区别于低层执行。与以往混淆推理与导航或提供丰富纠错反馈的基准不同,AsgardBench仅允许图像输入、动作历史及轻量级成功/失败信号,排除底层控制噪声,在受控模拟器中隔离交互式规划。基准包含108个任务实例,覆盖12类任务类型,通过对象状态、放置位置和场景配置的系统性变化,生成条件分支——同一指令可能需不同动作序列,强调执行中的条件分支与计划修复。对主流视觉语言模型的评估显示,缺少视觉输入时性能急剧下降,暴露出视觉接地与状态追踪的缺陷,最终影响交互式规划。本基准聚焦核心问题:模型能否真正利用所见信息,在预期之外时调整计划?
原文摘要 · Abstract (English)
With AsgardBench we aim to evaluate visually grounded, high-level action sequence generation and interactive planning, focusing specifically on plan adaptation during execution based on visual observations rather than navigation or low-level manipulation. In the landscape of embodied AI benchmarks, AsgardBench targets the capability category of interactive planning, which is more sophisticated than offline high-level planning as it requires agents to revise plans in response to environmental feedback, yet remains distinct from low-level execution. Unlike prior embodied AI benchmarks that conflate reasoning with navigation or provide rich corrective feedback that substitutes for perception, AsgardBench restricts agent input to images, action history, and lightweight success/failure signals, isolating interactive planning in a controlled simulator without low-level control noise. The benchmark contains 108 task instances spanning 12 task types, each systematically varied through object state, placement, and scene configuration. These controlled variations create conditional branches in which a single instruction can require different action sequences depending on what the agent observes, emphasizing conditional branching and plan repair during execution. Our evaluations of leading vision language models show that performance drops sharply without visual input, revealing weaknesses in visual grounding and state tracking that ultimately undermine interactive planning. Our benchmark zeroes in on a narrower question: can a model actually use what it sees to adapt a plan when things do not go as expected?
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。