arXiv:2602.06339cs.ROcs.AI2026-02被引 4

发现生成式机器人模型会幻觉出不物理的动作,影响可靠性。

Action Hallucination in Generative Vision-Language-Action Models

  • 分析生成式模型在动作生成中的拓扑、精度、时域三类结构性缺陷
  • 揭示动作幻觉源于可行行为与模型架构的不匹配,导致不可避的权衡
  • 为提升机器人策略可信度提供机制性改进方向,适合关注具身智能可靠性的研究者

生成式视觉-语言-动作模型(VLAs)有望实现端到端的通用机器人策略,但其能否真正解决具身环境下的动作生成根本问题仍不明确。本文通过分析违反物理约束的动作幻觉及其向规划层面的扩展,聚焦于潜在变量生成策略,揭示了可行机器人行为与常见模型架构之间的结构性错配。我们识别出三类关键障碍——拓扑、精度和时域,并证明它们会带来不可避免的权衡。该分析为已知的生成式机器人策略的实证失败提供了机制解释,并指出了在不牺牲表达能力的前提下提升可靠性和可信度的合理路径。

原文摘要 · Abstract (English)

Robot Foundation Models, such as VLAs, promise end-to-end generative robot policies with broad generalization. Yet it remains unclear whether they fundamentally resolve the core problem of action generation in embodied settings, or overcome the long-standing challenges of robotics. We address this question by analyzing action hallucinations that violate physical constraints and their extension to plan-level failures. Focusing on latent-variable generative policies, we show that hallucinations can arise from structural mismatches between feasible robot behavior and common model architectures. We study three such barriers -- topological, precision, and horizon -- and show how they impose unavoidable tradeoffs. Our analysis provides mechanistic explanations for reported empirical failures of generative robot policies and suggests principled directions for improving reliability and trustworthiness, without abandoning their expressive power.

机器人生成模型动作幻觉具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。