arXiv:2607.29169cs.ROcs.AI2026-07

让机器人视觉-动作保持同步,自动发现并修复运行时错误

ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency

论文配图:ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
图 1 · 摘自论文原文
  • 用动作条件的聚焦区域过滤无关视觉信息,保留关键接触与运动轨迹
  • 在局部遮挡下成功率从49.3%提升至90.3%,接近干净环境表现
  • 无需重训练即可应对视觉延迟、动作漂移等故障,适合部署于真实机器人

视觉-语言-动作(VLA)策略在机器人操作中表现优异,但易受运行时干扰影响,导致视觉观测、机器人状态与执行动作之间的时序对齐被破坏。本文提出ActFovea,一种即插即用的防护框架,可在不重新训练或修改原有VLA策略的前提下检测并缓解此类故障。ActFovea利用机器人运动学、本体感知状态和近期动作,构建动作条件的聚焦区域,保留接触相关区域和预测运动路径,同时抑制任务无关视觉内容。通过评估视觉运动与观测新鲜度是否与几何、本体感知和动作变化一致来检测运行风险。对于可恢复的干扰,生成特定扰动的候选观测,并在验证动作块后接受恢复;当观测过期或重放导致无法可靠恢复时,触发有限范围的安全失败机制。在多个LIBERO套件的闭环评估中,ActFovea将局部视觉遮挡下的成功率从49.3%提升至90.3%,弥补了93.7%的性能差距;在动作漂移和视觉延迟下分别提升7.0和9.8个百分点,且保持干净任务性能。在冻结观测重放测试中,所有试验均及时触发安全失败,无未防护失败案例。结果表明,时空视觉-动作一致性为VLA策略的运行时防护提供了有效基础。

原文摘要 · Abstract (English)

Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions. We introduce ActFovea, a plug-and-play safeguarding framework that detects and mitigates such failures without retraining or modifying the underlying VLA policy. ActFovea uses robot kinematics, proprioceptive states, and recent actions to construct action-conditioned foveated regions that retain contact-relevant areas and predicted motion corridors while suppressing task-irrelevant visual content. It detects runtime risks by evaluating whether visual motion and observation freshness remain consistent with geometric, proprioceptive, and action transitions. For recoverable disturbances, ActFovea constructs disturbance-specific candidate observations and accepts a recovery only after verifying the resulting action chunk. When stale or replayed observations make reliable recovery impossible, it invokes a bounded safe-failure procedure. In closed-loop evaluations of $π_0$ across multiple LIBERO suites, ActFovea increases success under localized visual overlays from 49.3\% to 90.3\%, closing 93.7\% of the gap to clean performance. It further improves success under action drift and visual delay by 7.0 and 9.8 percentage points, respectively, while preserving clean-task performance. Under frozen-observation replay, ActFovea triggers timely safe failure in all trials, with no unprotected failures. These results demonstrate that spatiotemporal visual-action consistency provides an effective basis for runtime safeguarding of VLA policies.

机器人视觉-动作鲁棒性防护机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。