arXiv:2601.09178cs.RO2026-01

用视觉提前预判环境变化,让机器人更早适应路况。

Vision-Conditioned Variational Bayesian Last Layer Dynamics Models

  • 用视觉信息动态调整车辆动力学模型,实现环境感知下的实时适应。
  • 在积水赛道上完成12圈全部成功,而无视觉的基线全数失控。
  • 适合高动态场景下的自动驾驶与机器人控制研究者参考。

敏捷的机器人控制需要提前预判环境对系统行为的影响。例如,驾驶员需感知前方道路以预估摩擦力并规划动作。在快速变化的条件下,实现自主框架中的主动适应仍具挑战。传统建模方法难以捕捉行为的突变,而自适应方法通常反应滞后,可能危及安全。本文提出一种视觉条件化的变分贝叶斯最后一层动态模型,利用视觉上下文来预测环境变化。模型先学习标准车辆动力学,再通过潜变量的特征仿射变换进行微调,实现上下文感知的动力学预测。该模型被集成至最优控制器用于赛车任务。我们在一辆雷克萨斯LC500上验证了该方法,使其在不同积水条件下完成全部12次尝试的赛道绕行。相比之下,所有无视觉上下文的基线方法均出现失控,证明了主动动态适应在高性能应用中的关键作用。

原文摘要 · Abstract (English)

Agile control of robotic systems often requires anticipating how the environment affects system behavior. For example, a driver must perceive the road ahead to anticipate available friction and plan actions accordingly. Achieving such proactive adaptation within autonomous frameworks remains a challenge, particularly under rapidly changing conditions. Traditional modeling approaches often struggle to capture abrupt variations in system behavior, while adaptive methods are inherently reactive and may adapt too late to ensure safety. We propose a vision-conditioned variational Bayesian last-layer dynamics model that leverages visual context to anticipate changes in the environment. The model first learns nominal vehicle dynamics and is then fine-tuned with feature-wise affine transformations of latent features, enabling context-aware dynamics prediction. The resulting model is integrated into an optimal controller for vehicle racing. We validate our method on a Lexus LC500 racing through water puddles. With vision-conditioning, the system completed all 12 attempted laps under varying conditions. In contrast, all baselines without visual context consistently lost control, demonstrating the importance of proactive dynamics adaptation in high-performance applications.

机器人控制视觉感知动态建模贝叶斯学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。