用自动化流程生成真实社交场景,评估机器人导航的社交能力。
SocRATES: Towards Automated Scenario-based Testing of Social Navigation Algorithms
- 用大模型自动生成符合场景和位置的交互任务
- 生成的场景能准确还原行人轨迹与行为,提升评估真实性
- 适合评估机器人社交导航算法,尤其对研究者和开发者有用
当前社交导航方法与基准主要关注人际距离和任务效率,但机器人社交能力的感知同样关键。本文提出基于场景的测试方法,通过具体人机交互场景揭示机器人行为特征。然而,手动构建此类场景成本高、难度大。为此,我们设计了一套自动化流水线,将简单场景元数据转化为详细文本场景,推断行人与机器人的轨迹,并模拟行人行为,实现更可控的评估。该流程利用大语言模型(LLMs)的社会推理与代码生成能力,简化场景生成与转换。实验表明,该方法生成的场景更真实,相比直接提示显著提升翻译质量。此外,我们还展示了专家可用性反馈及三类导航算法的案例评估结果。
原文摘要 · Abstract (English)
Current social navigation methods and benchmarks primarily focus on proxemics and task efficiency. While these factors are important, qualitative aspects such as perceptions of a robot's social competence are equally crucial for successful adoption and integration into human environments. We propose a more comprehensive evaluation of social navigation through scenario-based testing, where specific human-robot interaction scenarios can reveal key robot behaviors. However, creating such scenarios is often labor-intensive and complex. In this work, we address this challenge by introducing a pipeline that automates the generation of context-, and location-appropriate social navigation scenarios, ready for simulation. Our pipeline transforms simple scenario metadata into detailed textual scenarios, infers pedestrian and robot trajectories, and simulates pedestrian behaviors, which enables more controlled evaluation. We leverage the social reasoning and code-generation capabilities of Large Language Models (LLMs) to streamline scenario generation and translation. Our experiments show that our pipeline produces realistic scenarios and significantly improves scenario translation over naive LLM prompting. Additionally, we present initial feedback from a usability study with social navigation experts and a case-study demonstrating a scenario-based evaluation of three navigation algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。