用视觉模型识别潜水员动作,让水下机器人更好协作。
Semantically-Aware Diver Activity Recognition Framework for Effective Underwater Multi-Human-Robot Collaboration

- 基于注意力机制的框架,结合全局动作与局部交互语义
- 在2600+张标注图像上实现6类动作识别,性能优于现有模型
- 首个水下潜水员活动数据集,适合水下机器人研发者
在高风险、高挑战性的水下环境中,多智能体协同对拓展人类主导作业至关重要。为使自主水下航行器(AUV)真正成为可靠伙伴,必须具备环境理解与潜水员行为识别能力,以提供辅助并保障安全。为此,我们提出DAR-Net——一种基于Transformer的新型框架,用于分析复杂水下场景并分类潜水员活动。核心创新在于语义引导的学习范式,将基于Transformer的时序推理与像素级场景监督相结合,通过多损失训练策略,明确对齐全局活动识别与局部人机交互语义,尤其在低能见度条件下尤为重要。针对该领域数据稀缺问题,我们构建了首个水下潜水员活动(UDA)数据集,包含超过2,600张带像素级掩码的标注图像。在受控环境下严谨评估表明,DAR-Net在六类潜水员活动中实现了优异识别准确率,显著超越现有先进模型。尽管该数据集提供了关键基准,本工作仍属开创性探索,为未来更智能、更协同的水下机器人系统研究奠定基础。
原文摘要 · Abstract (English)
Effective multi-human-robot collaboration is essential for expanding human-led operations in the challenging and high-risk underwater environment. For autonomous underwater vehicles (AUVs) to become true teammates, they must be able to comprehend their surroundings and recognize a diver's activities to offer assistance and ensure safety. Towards this goal, we introduce DAR-Net, a novel transformer-based framework that analyzes complex underwater scenes to classify diver activities. Our contribution lies in a semantically guided learning formulation that couples transformer-based temporal reasoning with pixel-level scene supervision. This multi-loss training strategy explicitly aligns global activity recognition with local human-robot interaction semantics, which is particularly critical in low-visibility underwater conditions. To address the significant challenge of data scarcity in this domain, we present the first-ever Underwater Diver Activity (UDA) dataset, a foundational resource containing over 2,600 annotated images with pixel-level masks. Through rigorous experimental evaluations in a controlled environment, we demonstrate that DAR-Net achieves promising accuracy in recognizing six distinct diver activities, outperforming state-of-the-art models. While this dataset provides a crucial baseline, our work serves as a pioneering step, laying the groundwork for future research and facilitating the development of more intelligent, collaborative underwater robotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。