arXiv:2602.07413cs.RO2026-02

用动力系统建模机器人动作与视觉流的协同演化,实现高效自适应抓取。

Going with the Flow: Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity

  • 将技能建模为视觉与动作流的耦合动力系统,保证时序一致性。
  • 在7个仿真和4个真实任务中性能媲美或超越顶尖基线,推理速度更快。
  • 支持在线重规划,对遮挡鲁棒,适合高精度多指灵巧操作场景。

当前视觉-运动灵巧性模型通常依赖表达能力强的扩散和Transformer架构,但需大量数据与计算资源,且在多指灵巧操作上仍不可靠。这些模型将技能视为反应式映射,并依赖固定时长的动作分块,导致时序连贯性与反应性之间存在刚性权衡。为此,我们提出统一行为模型(UBM),将灵巧技能建模为耦合的动力系统,捕捉环境视觉特征(视觉流)与机器人本体状态(动作流)的协同演化。此类模型天然保证时序一致性,而非通过启发式平均。与预测任意动作影响的“世界模型”不同,UBM聚焦于编码已演示行为与期望环境影响之间的关系。一个UBM可视为伪规划器:给定初始条件,它能计算整个技能周期内的期望行为,同时‘想象’出对应的视觉特征流。为实现这一框架,我们提出Koopman-UBM(K-UBM),首个结构化潜在线性系统的实例。该模型计算高效,支持在线重规划:当预测与观测的视觉流偏差超过阈值时,模型自动触发重规划,作为自身运行时监控器。在7个仿真任务和4个真实任务中,我们的方法性能匹配或超越当前最优基线,同时具备更快推理、平滑执行、对遮挡鲁棒以及灵活重规划能力。

原文摘要 · Abstract (English)

Contemporary visuo-motor dexterity models often rely on expressive policy classes with diffusion and transformer backbones to achieve strong performance. However, these architectures require significant data and computational resources, and remain far from reliable, particularly for multi-fingered dexterity. Importantly, they model skills as reactive mappings and rely on fixed-horizon action chunking, creating a rigid trade-off between temporal coherence and reactivity. To address these issues, we first introduce Unified Behavioral Models (UBMs), a framework to represent dexterous skills as coupled dynamical systems that capture how visual features of the environment (visual flow) and proprioceptive states of the robot (action flow) co-evolve. As such, UBMs ensure temporal coherence by construction rather than heuristic averaging. Unlike world models that attempt to predict the impact of arbitrary robot actions on the environment, UBMs target behavioral dynamics that encode how demonstrated robot behavior is related to desired impacts on the environment. A UBM can be viewed as a pseudo planner: given an initial condition, it computes the desired robot behavior over the entire skill horizon, while simultaneously ``imagining" the resulting flow of visual features. To operationalize UBMs, we propose Koopman-UBM, a first instantiation of UBMs as a structured latent linear system. K-UBM is computationally efficient, enabling reactivity and adaptation via an online replanning strategy: the model acts as its own runtime monitor, automatically triggering replanning when predicted and observed visual flow diverge beyond a threshold. Across seven simulated tasks and four real-world tasks, our approach matches or exceeds the performance of state-of-the-art baselines, while offering considerably faster inference, smooth execution, robustness to occlusions, and flexible replanning.

灵巧操作动力系统在线重规划视觉-动作协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。