arXiv:2508.08983cs.ROcs.AI2025-08被引 1

通过推理动作背后的意图,用极少示范实现机器人任务泛化。

Rational Inverse Reasoning: Few-Shot Imitation by Inferring Intent through Planning

  • 从演示中推断隐藏的意图程序,将场景映射为可执行的任务规划。
  • 仅需一次示范即在新布局和新物体上成功执行,成功率提升28至34个百分点。
  • 适合需要少样本泛化的机器人学习研究者,尤其关注意图建模与规划结合。

人类仅需一两次示范即可学会新操作任务,并在新环境、新物体和新约束下完成。而现代机器人模仿学习通常需要数百至数千次示范,且在布局、几何或任务约束发生微小变化时性能显著下降。我们认为这一差距不仅源于数据量,更在于学习的抽象层次;真正的泛化需推断示范者行为背后的潜在意图,而非复制动作轨迹。本文提出理性逆推理(Rational Inverse Reasoning, RIR),将少样本模仿学习建模为对潜在解释程序的推断:紧凑且可执行的意图描述,能将以物体为中心的场景转化为结构化任务与运动规划(TAMP)的目标、子目标与约束。视觉语言模型生成候选程序,分层规划器提供有界理性似然。通过结合VLM程序提议与规划器反馈,RIR迭代优化候选集,逼近后验分布下的简洁可执行程序。在2D推理基准和真实Franka FR3机器人上,RIR仅需一次示范即可恢复可迁移的任务结构。在显著不同的布局与物体集上泛化,优于缺乏显式合理性与规划反馈的VLM-规划基线,在一示范与三示范设置下,下游成功率分别提升34和28个百分点。

原文摘要 · Abstract (English)

Humans can learn a new manipulation task from one or two demonstrations and then perform it in a new room, with new objects, under new constraints. Modern robot imitation learning, in contrast, typically needs hundreds to thousands of demonstrations and still degrades under modest shifts in layout, geometry, object set or task constraints. We argue this gap is not just about data, but also about the level of abstraction at which learning occurs; generalization requires inferring the latent intent underlying why a demonstrator behaved in a certain way, rather than reproducing how they moved. We present Rational Inverse Reasoning (RIR), which casts few-shot imitation as inference over latent explanation programs: compact, executable descriptions of intent that map an object-centric scene to a structured task-and-motion-planning (TAMP) specification of goals, subgoals and constraints. A vision-language model proposes candidate programs, and a hierarchical planner supplies a bounded-rational likelihood. By combining VLM program proposals, and planner-grounded feedback, RIR iteratively refines the candidate set to approximate a posterior over concise, executable programs. On a 2D reasoning benchmark and a real Franka FR3, RIR recovers transferable task structure from as little as one demonstration. Generalizing to substantially new layouts and object sets, RIR outperforms VLM-planning baselines that lack explicit rationality and planning-grounded inference, increasing downstream success rate by $34$ and $28$ percentage points in the one- and three-shot settings.

少样本学习意图推理机器人模仿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。