提出边界感知的完成检测机制,提升指令切换的稳定性和准确性。
Completion at the Boundary (CaB): Deployable Switching with Completion-Aware Control under Limited Calibration

- 用双侧边界信号构建完成判断,避免单变量失效
- 在Minecraft基准上实现92%复合指令成功率
- 适合部署于低校准、不可重训练的开放环境
视觉-语言-动作(VLA)智能体可执行自然语言指令,但实际部署系统仍缺乏可靠的完成判断接口:何时切换指令。这一问题在短序列复合指令(如“先做A,再做B”)中尤为突出,误判切换时机会引发下游失败连锁反应。完成判断本质上是闭环过程,因为切换行为会改变指令上下文,进而影响后续动作与观测。本文研究在可部署的低校准条件下(无测试时重训练、仅一次开发集校准的全局切换规则),如何实现可靠完成判断。在此约束下,将不对称边界证据压缩为单一标量易受任务极性变化影响。为此提出完成于边界(CaB)方法,通过保留事件局部的双侧边界证据,以边界阶段标记(Before/Hit/After)形式表示完成状态。CaB-When将此状态转换为最小化且可审计的切换决策(when),CaB-How则复用同一状态生成边界稳定的动作输出(how)。基于干预感知的E1/E2协议,在第一人称Minecraft VLA基准上验证,相同容量与部署条件下,CaB显著提升复合指令执行成功率与交接质量。
原文摘要 · Abstract (English)
Vision-language-action (VLA) agents can execute natural-language instructions, yet deployed systems still lack an operational interface: deciding when the instruction is complete. This gap is acute in short composites ("do A, then B"), where mistimed handoffs cascade into downstream failures. Completion is inherently closed-loop because switching is an intervention that changes the instruction context and thus future actions and observations. We study completion under a deployable low-calibration regime motivated by open-ended instruction spaces, enforcing no test-time relearning and a single globally calibrated switching rule selected once on development set and reused unchanged on test set. Under this constraint, collapsing asymmetric boundary evidence into a single scalar can be brittle under polarity shifts across tasks. We propose Completion at the Boundary (CaB), which predicts an event-local completion object in the form of Boundary-Phase Tokens (Before/Hit/After), retaining two-sided boundary evidence under this discipline. CaB-When converts this completion object into a minimal, auditable switching decision (when), while CaB-How reuses the same completion object to condition action generation for boundary-stable control through handoffs (how). Using an intervention-aware E1/E2 protocol, we show that CaB improves composite execution and handoff quality on a first-person Minecraft VLA benchmark under matched capacity and deployability constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。