arXiv:2606.05015cs.RO2026-06

通过无人机导航测试,发现预训练阶段模型鲁棒性决定真实世界泛化能力。

Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation

论文配图:Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation
图 1 · 摘自论文原文
  • 用视觉无人机在不同随机环境训练世界模型,跨环境验证其泛化性。
  • 预训练阶段表现好的模型全部成功实机飞行,最小过隙0.67米,仿真领先者失败。
  • 离散隐变量维度和训练序列长度是影响模型质量的关键因素。

世界模型是可预测环境演化的生成式模型,已成为提升机器人学习样本效率的有力工具。然而,其对环境变化的鲁棒性仍不明确。为此,我们以基于视觉的四旋翼无人机导航为测试任务,采用基于DreamerV3的世界模型,在不同环境随机性水平下进行训练,并通过跨环境验证评估所有水平下的性能,涵盖自监督学习(SSL)预训练与强化学习(RL)微调两个阶段。随后,我们将所有世界模型及其导航策略部署于真实四旋翼无人机,在未见过的环境中运行,包括一段开环测试:仅接收2.5秒真实感官输入后切断所有传感器,系统需完全依赖想象完成12米行进。结果表明,世界模型在SSL预训练阶段的鲁棒性是模拟到现实迁移成功的强预测因子:所有在跨环境SSL验证中表现良好的模型均在真实世界成功部署,穿过最小0.67米的间隙;而仿真表现最优的模型在真实平台失败。进一步分析发现,(a) 离散隐变量大小与 (b) 训练序列长度是主导世界模型质量的关键因素。

原文摘要 · Abstract (English)

World models, learned generative models that predict how an environment evolves, have become a promising tool for sample-efficient robot learning. Yet how robust they are to environmental variability remains poorly understood. To address this, we conduct a systematic study using vision-based quadrotor navigation as a testbed problem, training DreamerV3-based world models under varying levels of environmental randomness and evaluating them across all levels through cross-environment validation, spanning both Self-Supervised Learning (SSL) pretraining and Reinforcement Learning (RL) fine-tuning. We then deploy all world models and associated navigation policies on a real quadrotor in unseen environments, including an open-loop run where the model receives just 2.5s of real sensory input before all sensors are cut off, leaving the system to navigate entirely in imagination over a 12m traverse. Our results show that world model robustness during SSL pretraining is a strong predictor of sim-to-real transfer: every model that generalized well in cross-environment SSL validation deployed successfully in the real world, passing through gaps as narrow as 0.67m, whereas the model that dominated simulation policy evaluation failed on the real platform. We further identify (a) the discrete latent size and (b) the training-sequence length as the dominant factors governing world model quality.

世界模型四旋翼泛化性仿真实战

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。