让机器人任务评估更精准,结合动作、语言和视频变化判断进展
Action- and Language-Conditioned Video Assessment for Embodied Control

- 用视觉-语言模型分析动作与视频变化的匹配度
- 在模拟环境中误判率接近零,优于传统方法
- 适合需要精细反馈的机器人任务训练与优化
基于视觉的具身智能体执行多步骤自然语言指令时,需能评估完整轨迹的任务进展。传统方法依赖最终帧匹配或连续嵌入相似性,可能忽略中间关键过渡。我们提出ALVA(动作与语言条件化视频评估),一种基于视觉观察、执行动作序列和自然语言指令的轨迹评估方法。该方法分两阶段使用预训练视觉-语言模型:先以执行动作为条件,总结帧间视觉变化;再将生成摘要与指令对比,输出离散的轨迹级进展评分。在模拟3D家庭环境中,ALVA表现出保守评估模式,近零误报率。作为闭环策略优化的终端反馈,其效果优于静态图像和嵌入基视觉基线,显著缩小与真实标签基准的差距。结果表明,动作与语言条件化视频评估是可解释的反馈机制,适用于模拟具身控制任务。
原文摘要 · Abstract (English)
Vision-based embodied agents executing multi-step natural language instructions require feedback mechanisms that assess task progress over complete trajectories. Conventional approaches based on final-frame matching or continuous embedding similarity may overlook intermediate transitions that are necessary for determining whether an instruction has been completed. We propose ALVA (Action- and Language-Conditioned Video Assessment), a trajectory evaluator that conditions its assessment on visual observations, the executed action sequence, and the natural language instruction. The method uses a pre-trained vision-language model (VLM) in two stages: it first summarizes frame-to-frame visual transitions conditioned on the executed actions and then assesses the generated summary with respect to the instruction to produce a discrete trajectory-level progress score. In simulated 3D household environments, ALVA exhibits a conservative assessment pattern with near-zero false-positive rates. When used as terminal feedback for closed-loop policy optimization, it provides more effective feedback than the evaluated static image and embedding-based visual baselines and reduces the performance gap to a ground-truth oracle. These results support action- and language-conditioned video assessment as an interpretable feedback mechanism for the evaluated simulated embodied-control tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。