arXiv:2511.16050cs.RO2025-11

通过光照感知动作分块,提升水下机械臂在复杂光照下的操作稳定性

Bi-AQUA: Bilateral Control-Based Imitation Learning for Underwater Robot Arms via Lighting-Aware Action Chunking with Transformers

  • 采用双端控制与变压器分块动作策略,显式建模光照变化
  • 在多变光照下完成抓取、开抽屉、插销等任务,成功率超基线23%
  • 适合需力觉反馈的水下长时序高精度操作场景

水下机器人操作因光照变化、色彩衰减、散射和可视度下降而面临挑战。本文提出首个基于双端控制的模仿学习框架Bi-AQUA,其在策略中显式建模光照条件。Bi-AQUA融合基于Transformer的双端动作分块机制,采用分层光照感知设计,包括无标签光照编码器、基于FiLM的视觉特征调制以及用于动作条件化的光照标记。该设计使系统能在静态与动态变化的水下光照条件下自适应,同时保留双端控制的力觉敏感优势,尤其适用于长时序、接触密集型操作。真实世界实验在水下抓取、抽屉关闭和插销提取任务中表明,Bi-AQUA优于未建模光照的双端基线,在已见、未见及变化光照条件下均表现稳健。结果凸显了将显式光照建模与力觉感知的双端模仿学习结合对可靠水下操作的重要性。

原文摘要 · Abstract (English)

Underwater robotic manipulation remains challenging because lighting variation, color attenuation, scattering, and reduced visibility can severely degrade visuomotor policies. We present Bi-AQUA, the first underwater bilateral control-based imitation learning framework for robot arms that explicitly models lighting within the policy. Bi-AQUA integrates transformer-based bilateral action chunking with a hierarchical lighting-aware design composed of a label-free Lighting Encoder, FiLM-based visual feature modulation, and a lighting token for action conditioning. This design enables adaptation to static and dynamically changing underwater illumination while preserving the force-sensitive advantages of bilateral control, which are particularly important in long-horizon and contact-rich manipulation. Real-world experiments on underwater pick-and-place, drawer closing, and peg extraction tasks show that Bi-AQUA outperforms a bilateral baseline without lighting modeling and achieves robust performance under seen, unseen, and changing lighting conditions. These results highlight the importance of combining explicit lighting modeling with force-aware bilateral imitation learning for reliable underwater manipulation. For additional material, please check: https://mertcookimg.github.io/bi-aqua

水下机器人模仿学习光照建模双端控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。