arXiv:2607.10206cs.ROcs.AI2026-07

让流匹配模型可干预,实现多模态行为的主动选择

Source-Lifted Flow Matching for Intervenable Multimodal Imitation

论文配图:Source-Lifted Flow Matching for Intervenable Multimodal Imitation
图 1 · 摘自论文原文
  • 通过正交源提升机制,让不同行为路径由源点控制
  • 91.1%的干预能改变未来轨迹,消除路径交叉导致的混淆
  • 无需为每种模式单独建模,适合需要灵活行为选择的机器人任务

流匹配策略在模仿学习中表现优异,因其能建模复杂的多模态动作分布。然而其随机性是被动的:重复采样虽可产生多样行为,但用户无法从同一状态直接选择有效延续。本文提出源提升流匹配(SL-FM),一种可干预的流匹配策略,暴露可操作的源点控制变量,同时保持速度场共享且无需隐变量。核心机制为正交源提升,将不同控制源映射至辅助正交坐标,目标仍保留在原动作子空间,避免路径交叉歧义。通过端到端学习状态依赖的源混合,并设置责任下限,确保各控制变量持续有效,缓解死模式问题。在交叉流诊断与机器人控制基准测试中,SL-FM成功将被动源随机性转化为可操作干预变量,消除交叉引起的复合轨迹,91.1%的匹配前缀干预可改变未来路径,并在多个基准上实现显著性能提升。结果表明,仅通过源几何即可实现多模态行为的主动控制,无需对速度场按模式条件化。

原文摘要 · Abstract (English)

Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stochasticity is largely passive: repeated sampling may yield diverse behaviors, but users cannot directly choose among valid continuations from the same state. We propose Source-Lifted Flow Matching (SL-FM), a source-intervenable flow-matching policy that exposes such a handle while keeping the velocity field shared and latent-free. The handle selects only the source endpoint of the conditional flow, not a mode-specific field, preserving the standard formulation while avoiding decomposition into separate mode-conditioned dynamics. The core mechanism is \textbf{Orthogonal Source Lifting}, designed to prevent path-crossing ambiguity. Instead of partitioning target actions by mode, SL-FM lifts handle-specific sources into auxiliary orthogonal coordinates and keeps targets in the original action subspace. This preserves the demonstrated action distribution while allowing one shared field to carry different branches without merging at crossings. To keep handles usable across states, we learn a state-dependent source mixture end to end and use a responsibility floor, giving each handle weak supervision and mitigating dead modes. Experiments on crossing-flow diagnostics and robot-control benchmarks show that SL-FM converts passive source randomness into an actionable intervention variable. It removes crossing-induced composite trajectories, changes future routes in 91.1\% of matched-prefix interventions, and achieves strong free-deployment performance, with improvements in several benchmark settings. Overall, source geometry provides actionable multimodal control without conditioning the velocity field on the selected mode.

模仿学习多模态可干预流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。