arXiv:2607.29393cs.RO2026-07

提出多模态水下机器人动态预测模型,支持传感器失效时稳定运行

AquaJEPA: An Action-Conditioned Multimodal JEPA Family for Underwater Robot Dynamics

论文配图:AquaJEPA: An Action-Conditioned Multimodal JEPA Family for Underwater Robot Dynamics
图 1 · 摘自论文原文
  • 基于相机、声呐等多源数据,用动作条件建模预测未来状态
  • 在120个新场景中,成功率达状态仅方法提升12.5个百分点,误差降低0.189米
  • 可应对摄像头与声呐同时失效,适合水下作业高风险场景

水下机器人依赖互补传感器,其可靠性随能见度和运动状态急剧变化。本文提出AquaJEPA,一种可配置传感器的行动条件联合嵌入预测模型家族,覆盖全模态、仅相机、仅声呐及传感器丢失等多种配置。其成员共享潜在目标与滚动时域控制接口,从相机、前视声呐、本体感知和推进器指令中预测未来表征与物理动态。在石鱼(Stonefish)仿真环境中,使用一小时带动作标注的数据从头训练,评估涵盖120个全新配对场景,包含未见布局、能见度变化、动力学偏移及预定多普勒声纳(DVL)失效。AquaJEPA-base实现最强闭环性能,相比状态仅方法成功率提升12.5个百分点,最终误差减少0.189米;两者95%置信区间均不包含零。三组独立种子测试中,相比AquaJEPA-S,最终误差再降低0.118米,方向一致。AquaJEPA-robust在摄像头及摄像头-DVL双重失效期间,预测误差减少超一半。结果表明,全模态预测优于状态仅控制,且声呐单模态成员表现更优;传感器丢失训练显著提升传感器故障下的鲁棒性。

原文摘要 · Abstract (English)

Underwater robots rely on complementary sensors whose reliability changes abruptly with water visibility and vehicle motion. We introduce AquaJEPA, a sensor-configurable family of action-conditioned joint-embedding predictive models spanning full multimodal, camera-only, sonar-only, and sensor-dropout configurations. Its members share a latent objective and receding-horizon control interface that predict future representations and physical dynamics from camera, forward-looking sonar, proprioception, and thruster commands. Trained from scratch on one hour of action-labelled data, the family is evaluated in Stonefish on 120 fresh paired scenarios spanning unseen layouts, visibility changes, dynamics shifts, and scheduled DVL loss. AquaJEPA-base achieves the strongest aggregate closed-loop performance, improving success over state-only by 12.5 percentage points and reducing final error by 0.189 m; both paired 95% intervals exclude zero. In a separate three-seed evaluation, it reduces paired final error relative to AquaJEPA-S by 0.118 m, with the same direction for every seed. AquaJEPA-robust more than halves prediction error during camera and camera-DVL blackouts. These results show that full multimodal prediction improves over state-only control and the sonar-only family member in this benchmark, while sensor-dropout training provides robustness under sensor loss.

水下机器人多模态融合鲁棒控制预测模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。