arXiv:2510.02941cs.RO2025-10

对比量化指标与人工评估,找出行之有效的机器人导航评测方法

Metrics vs Surveys: An Analysis for Human-Aligned Benchmarking in Social Robot Navigation

  • 分析量化指标与人类主观评价的相关性,探索可替代调查的评测方式
  • 发现现有指标仅部分反映人类感知,关键主观因素仍被忽略
  • 适合需要快速评估机器人交互表现的研究者和开发者参考

社交导航(social navigation)是将移动机器人融入人类环境的关键挑战。其评估复杂,需兼顾舒适性、安全性与可理解性等多方面因素。当前主流的人类中心评估依赖问卷调查,虽可靠但成本高、难复现、难以跨系统比较。相比之下,数值化社会导航指标计算便捷,利于系统间对比,但学界尚未形成统一标准。本文探讨量化指标与人类评估之间的关联,旨在识别是否某些可量化的测量能反映人类感知。若存在相关性,这些指标可作为大规模调查不可行时的初步基准工具。结果表明,尽管现有指标捕捉了部分导航行为特征,但仍无法充分涵盖重要主观维度,亟需开发新指标以实现更符合人类偏好的评估体系。

原文摘要 · Abstract (English)

Social, also called human-aware, navigation is a key challenge for integrating mobile robots into human environments. The evaluation of such systems is complex, as factors such as comfort, safety, and legibility must be considered. Human-centered assessments, typically conducted through surveys, provide reliable insights but are costly, resource-intensive, and difficult to reproduce or compare across systems. Alternatively, numerical social navigation metrics are easy to compute and facilitate comparisons, yet the community lacks consensus on a standard set of metrics. This work explores the relationship between numerical metrics and human-centered evaluations to identify potential correlations. If specific quantitative measures align with human perceptions, they could serve as preliminary benchmarking tools, providing a human-aligned assessment when large-scale surveys are not feasible. Our results indicate that while current metrics capture some aspects of robot navigation behavior, important subjective factors remain insufficiently represented, necessitating new metrics.

机器人导航评估指标人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。