arXiv:2606.25504cs.RO2026-06中稿 · IROS 26

用自然语言生成逼真行人场景,让机器人练社交导航。

GROVE: Grounded Pedestrian Simulation via Natural Language for Interactive Social Robot Navigation

论文配图:GROVE: Grounded Pedestrian Simulation via Natural Language for Interactive Social Robot Navigation
图 1 · 摘自论文原文
  • 输入文字指令自动生成复杂行人行为场景。
  • 支持医院、办公室等多环境,模拟真实社交挑战。
  • 适配主流机器人仿真平台,提升训练真实感。

行人模拟是训练和部署社交机器人导航的关键,但现有系统高度僵化,需手动反复生成数据以定义简单场景。本文提出 GROVE,一种基于自然语言的文本到场景行人模拟框架,融合多种前沿方法,生成具有社会挑战性的高保真场景。用户可选择紧急、排队、正常等预设,或输入自定义提示词,实现高度定制化模拟。系统包含多个模块,分别保障长时序人类行为、中时序行人导航及短时序机器人社交互动的真实性和合理性。各模块根据提示动态调用最优模型,捕捉行人行为的情境细微差别,缩小仿真到现实(sim2real)的差距。行人模拟直接集成至 Isaac Sim、Gazebo 与 RViz 仿真器,支持机器人在高度社交环境中部署。我们在住宅、医院、办公等复杂场景下对 GROVE 与现有基线进行定性对比验证,结果表明其能生成多样化、复杂且逼真的行人行为,显著提升社交机器人导航的训练挑战性。

原文摘要 · Abstract (English)

Pedestrian simulation is a critical component for training and deploying social robot navigation approaches, yet it remains a largely rigid system that repeatedly requires manual data generation to define even simple scenarios. We propose GROVE, a text-to-scenario pedestrian simulation framework that combines state-of-the-art approaches to produce realistic, socially challenging scenarios for social robot navigation. Our framework allows users to customize one of several common presets (emergency, queuing, normal) or even enter a fully independent prompt to generate a highly customizable pedestrian simulation. Multiple modules separately ensure the realism and soundness of long-horizon human behavior, medium-horizon pedestrian navigation, and short-horizon robot/social interactions. Each module is tuned by the prompt in a way that reflects the user intent across all aspects of pedestrian simulation. By dynamically selecting one of several state-of-the-art (SotA) approaches in our modules based on the scenario, we capture many situational nuances of pedestrian behavior in order to narrow the simulation-to-real (sim2real) gap. The human simulation is directly integrated into Isaac Sim, Gazebo, and RViz simulators for robot deployment in highly social environments. We validate our approach through qualitative comparison against existing pedestrian simulation baselines across scenarios of varying complexity in residential, hospital, and office environments. The result is a high-fidelity pedestrian simulation that challenges social robot navigation with complex, diverse, realistic human behaviors.

社交机器人行人模拟自然语言生成仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。