让机器人通过预测进展、事件和不确定性来实现可解释的长程操作
LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation

- 引入三重预测机制:任务进展、语义事件、动作可靠性
- 在50个随机硬任务上达74.1%成功率,比基线高18.7个百分点
- 适合需要可解释性和鲁棒控制的通用机器人场景
通用视觉-语言-动作(VLA)策略主要依赖短时动作预测学习长程行为,但难以揭示采样命令之外的信息。这导致两个耦合瓶颈:单一动作目标需隐式包含任务进展、中间意图和局部可靠性,且这些控制状态在执行中不可见。受生物运动控制功能原理启发,我们提出LM-X,通过任务、事件和运动尺度上的显式预测组织控制,不追求解剖对应性。在线生成三种监督信号直接指导动作:返回值(RTG)衡量可见任务进展,事件剩余量(ETG)识别下一语义转换,异方差动作流通过传播方差估计局部可靠性。解释性因此内生于控制过程而非事后生成。在64块NVIDIA B200 GPU上进行20天预训练前,通过五任务控制门验证设计:完整模型相比仅动作基线提升16.0点成功率,相比最强单头变体提升10.8点。随后在超过20,000小时真实机器人轨迹(含1,000多小时失败回放)上训练,LM-X在50个随机硬任务上达成74.1%成功率,优于GR00T N1.7的55.4%;在7个真实任务上达68.6%,优于基线50.7%。RTG追踪语义进展与可视退化,方差在犹豫与振荡控制时上升。结果表明,显式多时标预测状态可增强控制并暴露可解释的内部估计。
原文摘要 · Abstract (English)
Generalist vision--language--action (VLA) policies learn long-horizon behavior mainly through short-horizon action prediction and reveal little beyond sampled commands. This creates two coupled bottlenecks: a single action target must implicitly absorb task progress, intermediate intent, and local reliability, while these control states remain hidden during execution. Inspired by functional principles of biological sensorimotor control, we introduce LM-X , which organizes prediction across task, event, and motor scales without claiming anatomical correspondence. Three explicitly supervised signals are emitted online and directly condition action generation: return-to-go (RTG) measures visible task progress, event-to-go (ETG) identifies the next semantic transition, and heteroscedastic action flow estimates local reliability through propagated variance. Explanation is therefore intrinsic to control rather than generated post hoc. Before a costly 20-day pretraining run on 64 NVIDIA B200 GPUs, a controlled five-task pretraining gate verifies the design: the complete model improves success by 16.0 points over the action-only backbone and by 10.8 points over the strongest single-head variant. We then train LM-X on more than 20,000 hours of real-robot trajectories, including over 1,000 hours of failed policy rollouts. LM-X achieves 74.1\% across 50 randomized-hard RoboTwin2.0 tasks versus 55.4\% for GR00T N1.7, and 68.6\% versus 50.7\% across seven real-robot tasks. RTG tracks semantic progress and visible regression, while variance rises during hesitation and oscillatory control. These results show that explicit multi-timescale predictive state can strengthen control while exposing interpretable internal estimates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。