用扩散模型统一生成可通行地图和路径,支持多类型机器人快速适配。
EmbodiedDiffusion: End-to-End Traversability-Guided Visual Diffusion for Heterogeneous Robot Navigation
- 基于扩散模型端到端预测可通行性并生成路径,无需独立规划器。
- 真实场景下导航成功率80%-100%,推理仅需90毫秒,新平台仅需10分钟数据微调。
- 通过模块化条件机制实现跨平台快速迁移,适合多机器人系统部署。
视觉可通行性估计是自主导航的核心,但现有方法要么依赖提示驱动的视觉-语言模型(VLM),要么将可通行性与路径规划分离,需独立规划器、大量建图、人工调参及长时间部署。我们提出EmbodiedDiffusion,一种基于扩散的框架,利用无规划器的合成监督和具身条件,从RGB图像中同时预测可通行性地图并生成可行路径,实现跨平台迁移。训练时,框架从一个大型VLM教师模型中提炼类别级可通行性语义至轻量学生模型,实现免提示、实时推理。采用基于FiLM的模块化条件机制,将具身特异性推理隔离为网络中可微调的小部分,使新机器人平台仅需10分钟视觉数据采集即可快速适应,无需重训视觉主干或轨迹扩散模型。在包含四足与空中机器人的室内环境中,EmbodiedDiffusion在全数据条件下实现80%-100%导航成功率,推理速度达90毫秒,展示了可扩展的统一可通行性推理与路径生成能力。
原文摘要 · Abstract (English)
Visual traversability estimation is central to autonomous navigation, yet most approaches either rely on prompt-driven Vision-Language Model (VLM) or decouple traversability from trajectory planning, requiring separate planners with heavy mapping, manual tuning, and extended deployment time. We propose EmbodiedDiffusion, a diffusion-based framework that simultaneously predicts traversability maps and generates feasible trajectories from RGB images using planner-free synthetic supervision and embodiment conditioning for cross-platform transfer. The framework distills category-level traversability semantics from a VLM teacher into a lightweight student model during training, enabling prompt-free, real-time inference at deployment. A modular FiLM-based conditioning mechanism isolates embodiment-specific reasoning into a compact trainable subset of the network, allowing rapid adaptation to new robot platforms without retraining the visual backbone or the trajectory diffusion model. Across indoor environments with quadruped and aerial robots, EmbodiedDiffusion achieves 80-100% navigation success in the full-data regime with real-time inference (90 ms) and adapts to new platforms using only 10 min of visual data collection, demonstrating scalable, unified traversability reasoning and trajectory generation for heterogeneous robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。