arXiv:2605.23856cs.RO2026-05被引 2

用点轨迹提升机器人动作模型,让系统更抗光照和遮挡干扰。

Point Tracking Improves World Action Models

论文配图:Point Tracking Improves World Action Models
图 1 · 摘自论文原文
  • 联合预测像素与2D点轨迹,显式建模运动变化。
  • 在长时序任务上显著优于纯像素模型,尤其在遮挡和离屏运动中。
  • 适合需要稳定运动感知的机器人控制任务。

机器人策略学习依赖于能捕捉环境动态的世界-动作模型,但基于像素的预测会将动态信息与光照、纹理等无关因素纠缠在一起,导致学习到的表征对任务无关的视觉变化敏感。本文提出JOPAT——一种联合像素与轨迹的世界-动作模型,通过一个去噪扩散Transformer同时预测潜在视觉观测、2D点轨迹及其可见性以及动作。核心思想是:点轨迹提供显式的运动表征,能捕捉长时序动态,在遮挡或部分离屏运动下依然稳健,比仅建模像素外观更具优势。在LIBERO和真实世界LeRobot任务上,JOPAT优于基于像素的基线模型,尤其在涉及遮挡、物体交互和离屏运动的长时序任务中提升最明显。

原文摘要 · Abstract (English)

Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and texture, making learned representations vulnerable to task-irrelevant visual variation. We propose JOPAT, a JOint Pixel-And-Track World-Action Model that predicts latent visual observations, 2D point tracks with visibility, and actions in a single denoising diffusion transformer. The key insight is that tracks provide an explicit representation of motion that captures long-horizon dynamics and remains robust under occlusion or partial out-of-frame motion, offering greater utility than modeling pixel appearance alone. On LIBERO and real-world LeRobot tasks, JOPAT improves over pixel-based baselines, with the largest gains on long-horizon tasks involving occlusion, object interaction, and off-screen motion.

机器人控制动作建模轨迹预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。