arXiv:2608.06898cs.ROcs.CL2026-08中稿 · the FoRMA workshop

为机器人选基础模型?这篇论文提出五维评估框架。

How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots

  • 构建三阶段评估流程:通用指标→模拟交互→真实机器人测试
  • 涵盖对话能力、用户安全等五大核心维度,覆盖全评估链路
  • 呼吁社区共建评估体系,推动机器人应用落地

研究者在为社交机器人选用基础模型时面临难题:如何选择?公开排行榜难以提供有效指导,因实时具身社交互动不在其关注范围内。直接实机评估又不切实际——每次实验需耗费大量人力、机器人和参与者时间。本文提出社交机器人基础模型的五个评估维度:(i) 对话能力,(ii) 用户安全,(iii) 具身角色表现,(iv) 场景适用性,(v) 受众适配性。为实现更高效、低成本的模型筛选,我们设计三阶段评估漏斗:先用通用指标初筛,再通过模拟交互深化评估,最后进行高成本的机器人专项测试。本文系统梳理了各维度在三个阶段的可用与缺失评估方法,并呼吁社区共同构建标准化评估框架。

原文摘要 · Abstract (English)

Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social interaction lie largely outside their focus. And direct evaluation is impractical at scale: each embodied study requires scarce participant, robot, and experimenter time. In this paper, we identify five evaluation dimensions for foundation models in social robots: (i) conversational competence, (ii) user safety, (iii) embodied character, (iv) target scene effectiveness, and (v) audience appropriateness. To make model selection cheaper and better informed, we propose a three-tiered evaluation funnel paradigm that first filters with general metrics, then extends to simulated interactions, and terminates in more expensive, robot-specific evaluation. We map all five dimensions across all three tiers, chart where applicable evaluation methods exist and are missing, and close with a call to action: let's build the evaluation framework together as a community.

机器人评估基础模型社交机器人评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。