arXiv:2608.13924cs.RO2026-08

解决机器人异步视觉-语言-动作控制中的行为不一致问题

BICPO-VLA: Behavior-Identified Continuation Preference Optimization for Smooth Asynchronous Vision-Language-Action Control

论文配图:BICPO-VLA: Behavior-Identified Continuation Preference Optimization for Smooth Asynchronous Vision-Language-Action Control
图 1 · 摘自论文原文
  • 通过因果历史编码器识别指令意图与任务进展
  • 分块生成动作并精确重构,减少迭代优化时间
  • 基于行为匹配的偏好优化,适配交接状态差异

请求到执行的过渡间隙由三重耦合因素造成:请求时刻行为意图模糊、动作生成期间物理状态漂移、新动作接管时的残余不兼容性。BICPO-VLA 依次应对:首先,使用指令感知的因果历史编码器识别命令所支持的行为及当前任务进度;其次,通过序列化哈爾子空间生成,将每个动作块分解为互补的成对支架与残差系数,实现两个专用生成阶段并完成精确重构,从而减少原始动作空间中的迭代优化,缩短机器人在新动作块可用前持续移动的时间间隔;最后,将已知的传出动作滚动至实际交接状态,对行为匹配候选者应用参考相对的Flow-DPO,适应剩余的请求-交接不匹配,且不改变其预期行为。

原文摘要 · Abstract (English)

The request-to-handoff gap has three coupled sources: ambiguity about the behavior intended at request time, physical-state drift accumulated during action generation, and residual incompatibility when the new action finally assumes control. BICPO-VLA addresses them in sequence. First, an instruction-aware causal history encoder identifies the behavior supported by the command and current task progress. Second, sequential Haar subspace generation decomposes each action chunk into complementary pairwise scaffold and residual coefficients, enabling two specialized generation stages followed by exact reconstruction. By reducing iterative refinement in the original action space, it shortens the interval over which the robot continues moving before the new chunk becomes available. Finally, BICPO rolls the known outgoing actions to the actual handoff state and applies reference-relative Flow-DPO among behaviorally matched candidates, adapting the generated chunk to the remaining request-to-handoff mismatch without changing its intended behavior.

机器人控制多模态动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。