arXiv:2605.09441cs.RO2026-05中稿 · RSS 2026

构建跨技能跨形态导航基准,推动通用机器人真实场景能力评估

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation

论文配图:Beyond Isolation: A Unified Benchmark for General-Purpose Navigation
图 1 · 摘自论文原文
  • 设计复合指令混合六类任务,强制智能体在单次任务中切换探索、交互与社交行为
  • 支持人形、四足、轮式机器人跨形态测试,覆盖170个融合真实扫描的环境
  • 采用人类远程操控生成1779条专家轨迹,捕捉真实行为细节,优于传统最短路径

通用具身智能体的发展受限于碎片化的评估协议,这些协议将导航技能孤立,并局限于特定机器人形态,无法反映真实世界中智能体需在不同形态上协调多种行为的复杂需求。为此,我们提出OmniNavBench,一个用于跨技能协同与跨形态泛化的统一基准。该基准实现三大范式转变:(1) 组合复杂性。提出复合指令,将点目标导航(PointNav)、视觉语言导航(VLN)、物体导航(ObjectNav)、社交导航(SocialNav)、跟人导航(Human Following)和问答导航(EQA)六类任务的子任务交错编排,迫使智能体在单个回合内完成探索、交互与社会合规的动态切换;(2) 形态普适性与传感器灵活性。构建无需依赖单一形态的仿真平台,支持人形、四足及轮式机器人跨形态泛化测试,配备模块化传感器接口,包含170个融合合成资产与真实扫描的环境;(3) 示范质量提升。摒弃仅依赖最短路径算法的轨迹生成方式,通过人类远程操控采集1779条专家轨迹,精准捕捉如探索性环视、预判规避等行为细节。大量实验表明,现有方法虽声称具备统一设计,却难以应对复杂交织的通用导航挑战。这揭示了当前能力与真实部署需求间的显著差距,凸显OmniNavBench作为新一代通用导航器测试平台的重要性。数据集、代码与排行榜已开放:http://omninavbench.cloud-ip.cc。

原文摘要 · Abstract (English)

The pursuit of general-purpose embodied agents is hindered by fragmented evaluation protocols that isolate navigation skills and fixate on specific robot morphologies, failing to reflect real-world scenarios where agents must orchestrate diverse behaviors across varying embodiments. To bridge this gap, we introduce OmniNavBench, a benchmark for cross-skill coordination and cross-embodiment generalization. OmniNavBench introduces three paradigm shifts: (1) Compositional Complexity. We propose composite instructions that interleave sub-tasks from 6 categories (PointNav, VLN, ObjectNav, SocialNav, Human Following and EQA), compelling agents to transition between exploration, interaction, and social compliance within a single episode. (2) Morphological Universality and Sensor Flexibility. We present a simulation platform that breaks the reliance on single-morphology evaluation, enabling generalization tests across humanoid, quadrupedal, and wheeled robots, with a modular sensor interface and 170 environments blending synthetic assets with real-world scans. (3) Demonstrations Quality. Moving beyond shortest-path algorithms, we curate 1779 expert trajectories via human teleoperation, capturing behavioral nuances such as exploratory glance and anticipatory avoidance. Extensive evaluations demonstrate that current methods, despite their claimed unified design, struggle with the complex, interleaved nature of general-purpose navigation. This exposes a critical disparity between existing capabilities and real-world deployment demands, underscoring OmniNavBench as a testbed for the next generation of generalist navigators. Dataset, code, and leaderboard are available at http://omninavbench.cloud-ip.cc.

通用导航多模态具身智能基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。