arXiv:2607.06522cs.AIcs.CV2026-07

让AI的推理与真实动作结果对齐,提升物理任务泛化能力

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment

论文配图:Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment
图 1 · 摘自论文原文
  • 设计双奖励机制,让AI推理贴合视觉环境和动作后果
  • 在未知任务和环境中,准确率显著高于基线模型
  • 适合研究视觉语言模型泛化与具身智能的学者

视觉语言模型(VLMs)在交互式物理推理中难以泛化,尤其在未见过的任务和环境中。主要问题包括:与物理现实矛盾的虚构思维链(CoT),以及推理与行为之间的不一致。我们提出VAORA(视觉动作结果推理对齐)方法,通过两个互补奖励解决:视觉对齐奖励将模型推理锚定在视觉上下文,不受具体动作影响;视觉-动作对齐奖励则将推理与模型动作所引发的视觉结果绑定。结合预训练领域专家代理估计的成功概率,采用平滑密集奖励以增强训练稳定性。在PHYRE和Virtual Tool数据集上的实验表明,该方法在新任务和未知环境设置下表现优异,验证了通过VAORA可有效引导出具有一致性且可泛化的物理智能。

原文摘要 · Abstract (English)

Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. Two key failure modes are prominent: hallucinated chain-of-thought (CoT) reasoning that contradicts physical reality, and misalignment between the model's reasoning and actions. We present VAORA (Visual Action Outcome Reasoning Alignment), a novel reward design that directly addresses both issues. VAORA introduces two complementary rewards: Visual Alignment Reward, which anchors VLM reasoning to the visual context independent of the agent action itself, and Visual-Action Alignment Reward, which grounds reasoning in the visual outcome induced by the model's action. Together, these rewards suppress hallucinated CoT and reduce the gap between reasoning and behavior. To improve training stability, we further employ smooth, dense rewards by estimating success probabilities using a pre-trained in-domain expert agent. Experiments on PHYRE and Virtual Tool support our performances across novel-task and unseen-environment settings, confirming that grounded and generalizable physical intelligence can be induced through VAORA.

物理推理视觉语言模型具身智能强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。