首个水下具身智能体基准,推动AI在复杂海洋环境中的自主探索。
OceanGym: A Benchmark Environment for Underwater Embodied Agents
- 基于多模态大模型构建统一框架,融合感知、记忆与决策
- 涵盖8个真实场景任务,验证了当前模型与人类专家的显著差距
- 适合研究水下机器人、具身智能与多模态学习的学者使用
我们提出OceanGym,首个面向水下具身智能体的综合性基准平台,旨在推动人工智能在最具挑战性的现实环境之一——海洋中的发展。与陆地或空中环境不同,水下场景面临低能见度、动态洋流等极端感知与决策难题,导致有效智能体部署极为困难。OceanGym包含八个高保真任务域,采用由多模态大语言模型(MLLM)驱动的统一智能体框架,整合感知、记忆与序列决策能力。智能体需理解光学与声呐数据,在复杂环境中自主探索,并在恶劣条件下完成长时程目标。大量实验表明,当前最先进的MLLM驱动智能体与人类专家之间存在显著性能差距,凸显了水下环境下感知、规划与适应能力的持续挑战。通过提供高保真、严谨设计的测试平台,OceanGym为开发鲁棒的具身智能提供了基础,并推动其向真实水下无人系统迁移,标志着迈向能在地球最后未探索前沿自主运行智能体的重要一步。代码与数据已公开于https://github.com/OceanGPT/OceanGym。
原文摘要 · Abstract (English)
We introduce OceanGym, the first comprehensive benchmark for ocean underwater embodied agents, designed to advance AI in one of the most demanding real-world environments. Unlike terrestrial or aerial domains, underwater settings present extreme perceptual and decision-making challenges, including low visibility, dynamic ocean currents, making effective agent deployment exceptionally difficult. OceanGym encompasses eight realistic task domains and a unified agent framework driven by Multi-modal Large Language Models (MLLMs), which integrates perception, memory, and sequential decision-making. Agents are required to comprehend optical and sonar data, autonomously explore complex environments, and accomplish long-horizon objectives under these harsh conditions. Extensive experiments reveal substantial gaps between state-of-the-art MLLM-driven agents and human experts, highlighting the persistent difficulty of perception, planning, and adaptability in ocean underwater environments. By providing a high-fidelity, rigorously designed platform, OceanGym establishes a testbed for developing robust embodied AI and transferring these capabilities to real-world autonomous ocean underwater vehicles, marking a decisive step toward intelligent agents capable of operating in one of Earth's last unexplored frontiers. The code and data are available at https://github.com/OceanGPT/OceanGym.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。