首个融合物理与经济数据的导航评估基准,衡量机器人商业可行性。
CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents
- 用真实财务与监管数据评估机器人导航的盈亏表现
- 7种方法均亏损,最高仅实现-28.40美元/次的最小亏损
- 支持真实机器人验证,推动算法向商业化落地转化
现有导航评估侧重任务成功率,却忽略商业部署必需的经济约束。我们提出CostNav——首个基于物理模拟的经济导航基准,结合Isaac Sim的碰撞与载货动力学,以及证券交易委员会(SEC)文件和简略伤害等级(AIS)报告等产业标准数据,量化导航指标与商业部署间的差距。结果显示,高成功率并不等同于经济可行。评估七种基线(两规则、五模仿学习),所有方法均不盈利,贡献毛利为负。仅使用RGB相机和GPS的CANVAS在达标率(SLA)非零的方法中表现最佳,单位运行成本-28.40美元,优于配备激光雷达的Nav2+GPS(-37.34美元/次)。在真实配送机器人上测试的模拟训练策略,其SLA达标率接近仿真结果,表明模型性能可跨仿真与现实迁移。我们呼吁社区挑战在CostNav上实现经济可行性,该基准以成本收益结果评分。所有资源已开源:https://github.com/worv-ai/CostNav。
原文摘要 · Abstract (English)
Current navigation benchmarks focus on task success but do not capture the economic constraints essential for commercializing autonomous delivery systems. We introduce CostNav, an Economic Navigation Benchmark that evaluates physical AI agents on a cost-revenue and break-even analysis, pairing Isaac Sim's collision and cargo dynamics with industry-standard data such as Securities and Exchange Commission (SEC) filings and Abbreviated Injury Scale (AIS) injury reports. To our knowledge, CostNav is the first physics-grounded economic benchmark to use regulatory and financial data to quantify the gap between navigation metrics and commercial deployment, revealing that high task-success rates alone do not ensure economic viability. Evaluating seven baselines (two rule-based and five imitation-learning methods), we find no method economically viable: all yield negative contribution margins. CANVAS, using only an RGB camera and GPS, attains the highest task success and the least-negative margin among methods with non-zero Service-Level Agreement (SLA) compliance (-\$28.40/run), outperforming LiDAR-equipped Nav2 w/ GPS (-\$37.34/run). A sim-trained policy evaluated on a real delivery robot yields SLA compliance close to its simulation result, indicating that policy performance in CostNav's simulation transfers to real-world deployment. We challenge the community to achieve economic viability on CostNav, which scores methods by cost-revenue outcomes. All resources are available at https://github.com/worv-ai/CostNav.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。