arXiv:2603.26322cs.RO2026-03被引 1

用视觉直接生成导航与抓取动作,5分钟自监督数据即可上手。

DiffusionAnything: End-to-End In-context Diffusion Learning for Unified Navigation and Pre-Grasp Motion

  • 通过多尺度特征调制统一处理远距离导航和近距离操作
  • 零样本泛化到新场景,仅需RGB输入,10Hz推理速度
  • 无需语言模型或深度传感器,适合嵌入式部署

高效地从视觉直接预测运动规划仍是机器人领域的核心挑战,传统方法需明确目标设定和任务定制。近期视觉-语言-动作(VLA)模型虽可直接从视觉输入推断动作,但需大量计算资源、训练数据,且在新场景中无法零样本泛化。本文提出一种统一的图像空间扩散策略,通过多尺度特征调制实现米级导航与厘米级操作一体化,每任务仅需5分钟自监督数据。三大创新:(1)多尺度FiLM条件控制任务模式、深度尺度与空间注意力,使单一模型适配不同行为;(2)轨迹对齐的深度预测聚焦生成路径上的度量三维推理;(3)来自AnyTraverse的自监督注意力实现无语言模型、无深度传感器的目标导向推断。系统仅依赖RGB输入(2.0 GB内存,10 Hz),实现对新场景的鲁棒零样本泛化,同时满足机载部署需求。

原文摘要 · Abstract (English)

Efficiently predicting motion plans directly from vision remains a fundamental challenge in robotics, where planning typically requires explicit goal specification and task-specific design. Recent vision-language-action (VLA) models infer actions directly from visual input but demand massive computational resources, extensive training data, and fail zero-shot in novel scenes. We present a unified image-space diffusion policy handling both meter-scale navigation and centimeter-scale manipulation via multi-scale feature modulation, with only 5 minutes of self-supervised data per task. Three key innovations drive the framework: (1) Multi-scale FiLM conditioning on task mode, depth scale, and spatial attention enables task-appropriate behavior in a single model; (2) trajectory-aligned depth prediction focuses metric 3D reasoning along generated waypoints; (3) self-supervised attention from AnyTraverse enables goal-directed inference without vision-language models and depth sensors. Operating purely from RGB input (2.0 GB memory, 10 Hz), the model achieves robust zero-shot generalization to novel scenes while remaining suitable for onboard deployment.

扩散模型机器人视觉导航自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。