用人类视频训练机器人完成复杂长任务,成功率超74%。
Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation

- 用实例感知的三元组表示高层目标,实现任务分解与跟踪
- 在真实场景中达74.8%成功率,支持对象级和组合泛化
- 闭环系统在线监测进展,避免执行偏差和幻觉
在不确定环境中保障长时程机器人操作的可靠性仍具挑战。端到端视觉语言动作(VLA)模型依赖大量数据且不可解释,难以诊断验证;分层流水线虽更可解释,但其规划缺乏观测支撑,与底层动作对齐弱,且无在线反馈,导致开环行为和幻觉。为此,我们提出三元组到轨迹系统(TTS),一种基于人类视频的闭环长时程模仿学习系统,减少对机器人采集数据的依赖。TTS将高层子目标表示为实例感知的三元组,将其转换为连续轨迹先验以执行,并从观测中监控任务进展以实现在线重规划。在多种真实世界长时程任务中,TTS平均成功率达74.8%,支持对象级与组合泛化。
原文摘要 · Abstract (English)
Ensuring reliability in uncertain environments remains difficult for long-horizon robotic manipulation. End-to-end VLA models are data-heavy and opaque, making diagnosis and verification difficult. Hierarchical pipelines are more interpretable, but their plans are often weakly grounded in observations, weakly aligned with low-level actions, and computed without online feedback, leading to open-loop behavior and hallucinations. To address these issues, we introduce the Triplet-to-Track System (TTS), a closed-loop long-horizon imitation learning system that uses human videos to reduce reliance on robot-collected data. TTS represents high-level subgoals as instance-grounded triplets, translates them into continuous track priors for execution, and monitors task progress from observations for online replanning. Across diverse real-world long-horizon tasks, TTS achieves a 74.8\% average success rate and supports object-level and compositional generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。