通过发现失败边界提升视觉语言动作模型的鲁棒性
Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models

- 基于真实与模拟协同训练,用数字孪生探测失败边界
- 利用进度感知信号定位失败发生的具体阶段
- 无需额外奖励函数,直接优化动作流方向以增强容错能力
视觉-语言-动作(VLA)模型经监督微调后存在结构不对称:专家示范仅告知成功行为的位置,却未提供其失效范围。我们提出应将鲁棒适应视为‘失败边界学习’——即发现、定位并塑造可恢复偏差与任务失败之间的边界。为此,我们提出DLS框架,基于少量真实示范和模拟协同训练构建真实行为先验,通过在线策略数字孪生滚动推演大规模发现失败边界。不同于将每条轨迹简化为二值标签,语义进展定位利用特权模拟状态,赋予具有进展感知的信号,精确捕捉失败边界被跨越的时刻而非仅判断是否失败。这些信号驱动动作流中的方向性边界塑形:强化产生成功的去噪方向,抑制导致失败的方向,无需动作概率或辅助评判器。在真实机器人操作任务中,DLS在随机初始状态和未见视觉条件下均显著优于监督微调与在线强化学习基线。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models adapted through supervised fine-tuning (SFT) inherit a structural asymmetry: expert demonstrations teach the policy where success behavior lies, but provide no signal about where it ceases to be reliable. We argue that robust VLA adaptation should therefore be viewed not as further demonstration fitting, but as **Failure-Boundary Learning**---the problem of *Discovering*, *Localizing*, and *Shaping* the boundary between recoverable deviations and task failure. To instantiate this view, we propose **DLS**: built on a **real-grounded behavioral prior** from few real demonstrations and simulated co-training, DLS *discovers* failure boundaries at scale through on-policy digital twin rollouts. Rather than reducing each rollout to a binary label, **semantic progress localization** uses privileged simulator states to assign progress-aware signals that capture *where* the failure boundary is crossed, not merely *whether*. These signals drive **directional boundary shaping** in the flow dynamics---reinforcing success-producing denoising directions and suppressing failure-producing ones, without action likelihoods or auxiliary critics. Across real-robot manipulation tasks, DLS improves robustness over SFT and online RL baselines, especially under randomized initial states and unseen visual conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。