arXiv:2601.07060cs.RO2026-01被引 25

让机器人做复杂任务时能看懂进度、识别关键操作点,减少重复出错。

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

  • 基于物体交互线索和子任务进度,构建可指导动作的视觉推理机制
  • 在真实场景中实现91.8%成功率,比基线提升12.5%以上
  • 适合需要连续多步操作的机器人智能控制研究者

视觉-语言-动作(VLA)模型在机器人操作中展现出潜力,但在长时序、多步骤任务上仍表现不佳。现有方法缺乏内部推理机制,无法识别任务相关交互线索或追踪子任务进展,导致重复动作、漏步和过早终止等严重错误。为此,我们提出PALM框架,以交互为中心的可感知性推理与子任务进度提示为核心,构建策略学习。PALM提取融合物体相关性、接触几何、空间位置和运动动力学的互补可感知表征,作为视觉-运动控制的任务锚点。为增强长周期执行稳定性,该框架预测连续子任务内进度,实现平滑过渡。在大量仿真与真实世界实验中,PALM持续优于基线,在LIBERO-LONG数据集上达91.8%成功率,在CALVIN ABC->D上平均任务长度提升12.5%,并在三个长时序泛化设置中实现实验室基准2倍性能提升。

原文摘要 · Abstract (English)

Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify task-relevant interaction cues or track progress within a subtask, leading to critical execution errors such as repeated actions, missed steps, and premature termination. To address these challenges, we introduce PALM, a VLA framework that structures policy learning around interaction-centric affordance reasoning and subtask progress cues. PALM distills complementary affordance representations that capture object relevance, contact geometry, spatial placements, and motion dynamics, and serve as task-relevant anchors for visuomotor control. To further stabilize long-horizon execution, PALM predicts continuous within-subtask progress, enabling seamless subtask transitions. Across extensive simulation and real-world experiments, PALM consistently outperforms baselines, achieving a 91.8% success rate on LIBERO-LONG, a 12.5% improvement in average length on CALVIN ABC->D, and a 2x improvement over real-world baselines across three long-horizon generalization settings.

机器人操控长程任务视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。