arXiv:2607.25219cs.RO2026-07

构建逼真3D视觉仿真平台,推动机器人社交导航研究

SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation

论文配图:SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation
图 1 · 摘自论文原文
  • 用3D高斯泼溅技术建模场景与行人,实现真实感视觉输入
  • 引入大模型生成语义轨迹,结合条件生成器实现自然连续动作
  • 提供分难度评估集与多维度评测指标,适合视觉导航算法验证

社交导航正从简化2D环境转向基于视觉的通用场景,要求机器人仅凭机载视觉观察实现社会合规行为。然而现有仿真平台或缺乏视觉输入,或缺少移动人类角色,或在外观与行人行为上缺乏真实感,难以支撑视觉导航发展。本文提出SONG平台,基于3D高斯泼溅(3DGS)同时表现场景与行人,利用大语言模型生成语义轨迹,并通过轨迹条件生成器合成全身运动,实现连续自然的行人行为。在此基础上,我们构建SONG-Bench评估集,按难度分层设计测试用例,并提出涵盖有效性、安全性与社会合规性的多维评价体系。对代表性导航基线的系统评估发现:(a) 视觉社交导航尚未解决;(b) 安全性缺陷先于社交礼仪问题;(c) 真实世界数据比模型规模更重要。关键的是,我们在自建数据上微调显著提升了真实环境中的成功率。期望该平台为下一代视觉社交导航研究提供可信严谨的测试基准。

原文摘要 · Abstract (English)

Social navigation has progressed from simplified 2D environments toward a more general vision-based setting, in which a robot needs to achieve socially compliant behavior purely from onboard visual observations. Yet supporting simulation platforms have not kept pace: existing options either lack visual observations, lack moving human avatars, or fall short of real-world fidelity in appearance and pedestrian behavior, offering limited support for advancing vision-based social navigation. We introduce SONG, a SOcial Navigation platform powered by 3D Gaussian splatting (3DGS). It leverages 3DGS for both scene and avatar representations, drives pedestrians using semantically grounded trajectories generated by a large language model, and synthesizes their full-body motion with a trajectory-conditioned generator to produce continuous, natural movement. On top of the platform, we curate SONG-Bench, a set of evaluation episodes stratified by difficulty, and propose a multi-dimensional metric suite covering effectiveness, safety, and social compliance. A systematic evaluation of representative navigation baselines reveals three findings: (a) vision-based social navigation is far from solved; (b) a critical safety deficit precedes social etiquette; (c) real-world data matters more than model scale. Crucially, we demonstrate that fine-tuning on our curated data effectively improves the success rate in real-world environments. We hope our platform provides a faithful and rigorous testbed for the next generation of vision-based social navigation research.

社交导航3D高斯仿真平台视觉导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。