arXiv:2511.13312cs.ROcs.AI2025-11

用扩散模型生成语言控制的机器人操作轨迹,提升多任务执行成功率。

EL3DD: Extended Latent 3D Diffusion for Language Conditioned Multitask Manipulation

  • 结合视觉与文本输入,通过扩散模型生成精准机器人动作序列。
  • 在CALVIN数据集上实现多任务连续执行成功率提升,长周期任务表现更优。
  • 适合关注语言指令控制机器人、具身智能的研究者参考。

在人类环境中行动是通用机器人的重要能力,需要对自然语言有稳健理解并应用于物理任务。本文将扩散模型融入视觉-运动策略框架,融合视觉与文本输入,生成精确的机器人轨迹。训练时利用参考演示,使模型能根据文本指令在机器人邻近环境中执行操作任务。研究通过改进嵌入表示,并借鉴图像生成中的扩散技术,扩展了现有模型。在CALVIN数据集上的评估表明,该方法在多种操作任务中表现更优,多任务顺序执行时的长周期成功率显著提高。本方法验证了扩散模型在具身智能中的有效性,推动了通用多任务操作的发展。

原文摘要 · Abstract (English)

Acting in human environments is a crucial capability for general-purpose robots, necessitating a robust understanding of natural language and its application to physical tasks. This paper seeks to harness the capabilities of diffusion models within a visuomotor policy framework that merges visual and textual inputs to generate precise robotic trajectories. By employing reference demonstrations during training, the model learns to execute manipulation tasks specified through textual commands within the robot's immediate environment. The proposed research aims to extend an existing model by leveraging improved embeddings, and adapting techniques from diffusion models for image generation. We evaluate our methods on the CALVIN dataset, proving enhanced performance on various manipulation tasks and an increased long-horizon success rate when multiple tasks are executed in sequence. Our approach reinforces the usefulness of diffusion models and contributes towards general multitask manipulation.

机器人扩散模型语言控制多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。