用粒子分解机器人动作与外观,提升视频预测精度。
AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction

- 将机器人视频分解为可追踪的语义粒子,分治运动与外观建模。
- 相比基线,轨迹误差降低21.0%,未来帧PSNR提升至20.573。
- 适合需要精准动作推理的机器人视觉任务,对语言理解要求高。
机器人视频预测需兼顾精确运动推理与高频外观保持,但传统像素模型常将二者纠缠,且被强基线掩盖真实进展。我们提出AcrossVAM1.0,一种轻量级、文本辅助的视频动作模型,将未来预测解耦为基于物体的运动与稠密外观建模。采用冻结的SAM3-DLP编解码器,将四帧上下文分解为机器人、手臂、夹爪及背景的语义粒子。一个仅0.28M参数的时空Transformer对齐粒子身份,前推状态,并通过冻结的OpenCLIP指令嵌入进行FiLM调制。因果双流解码器结合粒子渲染运动与仅来自最后一帧的外观编码;残差修正器与学习到的投送掩码生成五帧未来图像,无需访问未来外观。在基于多样真实机器人轨迹构建的VRS基准上,粒子动态使轨迹误差降低21.0%。在三个投送掩码种子下,未来帧PSNR/SSIM从19.97/0.796提升至20.573/0.8004;粒子生成本身使运动区域PSNR从11.89提升至13.23。当前模型尚未超越持久性基线在LPIPS上的表现,正确与打乱语言对轨迹误差影响仅2.8–3.1%。我们报告这些局限性,并附带最优、负控、多种子及按机器人分析结果。结果表明,显式粒子动态是机器人视频预测中一个有前景的低维接口,而鲁棒的语言接地与外观投送仍是主要挑战。
原文摘要 · Abstract (English)
Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet monolithic pixel models entangle these objectives and often conceal their progress behind a strong last-frame baseline. We present AcrossVAM1.0, a lightweight, text-assisted video action model that factorizes future prediction into object-centric motion and dense appearance. A frozen SAM3-DLP codec decomposes four context frames into semantic particles for the robot, arm, and gripper, together with a background latent. A 0.28M-parameter spatio-temporal Transformer aligns particle identities, rolls their states forward, and is modulated by a frozen OpenCLIP instruction embedding through FiLM. A causal dual-stream decoder combines particle-rendered motion with appearance encoded exclusively from the last observed frame; a residual refiner and learned delivery mask produce five future frames without access to future appearance. On our VRS benchmark constructed from diverse real-robot trajectories, particle dynamics reduce trajectory error by 21.0\% over persistence. Across three delivery-mask seeds, AcrossVAM1.0 improves future-frame PSNR/SSIM from 19.97/0.796 to 20.573/0.8004, while raw particle generation improves motion-region PSNR from 11.89 to 13.23. The delivered model does not yet beat persistence in LPIPS, and correct-versus- shuffled language changes trajectory error by only 2.8--3.1%. We report these limitations alongside oracle, negative-control, multi-seed, and per-robot analyses. The results show that explicit particle dynamics are a promising low-dimensional interface for robot video prediction, while robust language grounding and appearance delivery remain the principal open challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。