arXiv:2607.10781cs.CVcs.RO2026-07中稿 · Robotics: Science …

不重训练,仅靠采样时注入能量函数即可控制驾驶模型行为。

Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models

论文配图:Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models
图 1 · 摘自论文原文
  • 采样时用可微能量函数引导轨迹生成,无需修改模型权重。
  • 在Open-Sora 2.0基础上实现反事实刹车,轨迹与视频同步对齐。
  • 发现跨流耦合是实现端到端可控推演的关键机制,适合自动驾驶研究者。

基于大型视频扩散模型的驾驶世界模型虽能生成逼真场景,但难以控制:通常需重训练或手工布局条件才能遵循交通规范。我们探究是否真的需要训练来实现可控性。实验表明,一种联合生成未来视频与规划自车轨迹的修正流驱动模型,可在采样阶段仅通过编码交通规范的可微能量函数,完全实现轨迹引导,无需对扩散主干进行知识特异性重训练。具体而言,基于Open-Sora 2.0 MM-DiT主干的模型,可通过采样时注入能量指导实现反事实刹车。然而我们发现,生成视频尚未随引导轨迹通过主干的联合自注意力机制对齐,揭示跨流耦合是实现端到端可控推演的关键要求。

原文摘要 · Abstract (English)

Driving world models built on large video-diffusion backbones generate realistic scenes but are hard to control: enforcing a traffic norm typically means retraining the backbone or conditioning it on hand-built layouts. We ask whether controllability requires training at all. Our experiment shows that a rectified-flow driving world model, which jointly generates future video and a planned ego trajectory, can have its planned trajectory steered entirely at sampling time by differentiable energy functions that encode driving norms, without knowledge-specific retraining of the diffusion backbone. Concretely, we demonstrate that a world model built on Open-Sora 2.0 MM-DiT backbone can be steered to brake at a counterfactual target by injecting energy guidance at sampling time. However, we find that the generated video does not yet follow the steered trajectory through the backbone's joint self-attention and identify the cross-stream coupling as a crucial requirement for end-to-end-controllable rollouts.

驾驶模型扩散模型能量引导无训练控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。