arXiv:2605.31119cs.ROcs.LG2026-05中稿 · 2026 IEEE/RSJ Inte…

让机器人从意外中学习,避免重复踩坑。

Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning

论文配图:Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning
图 1 · 摘自论文原文
  • 通过视觉语言模型分析异常原因,结合上下文推断风险
  • 用核回归实现少量样本下对突发问题的快速建模
  • 适合在复杂真实环境中长期运行的自主机器人

在机器人领域,危险和困境往往具有实体特异性且相对特定代理。当前前沿目标是使自主移动机器人能在未见过的非结构化环境中有效运作。一个关键挑战在于无法预知所有针对特定机器人的潜在危险。尽管近期研究利用大型基础视觉-语言模型(VLM)预先预测常见常识性风险,但仍难以捕捉可能的交互与实体依赖性困境。本文提出一种持续学习框架,使移动具身智能体能够在线从干扰中学习,并通过语义归因将异常行为关联至原因,从而提升未来对世界的预测与规划能力。该框架名为“Don't Fool Me Twice”,首先观测干扰并描述其对机器人的影响;此描述结合视觉上下文,通过VLM查询可能成因;局部干扰采用核回归表征,实现对瞬态异常的高效、少样本建模。我们采用语义体素中心建模估计认知不确定性,使交互驱动的干扰被视作可学习的空间行为,支持更丰富的后续恢复策略。我们在仿真与硬件上验证了四个假设,覆盖多种机器人形态与困境类型。

原文摘要 · Abstract (English)

In robotics, dangers and adversity modes are often embodiment-specific and relative to each agent. A frontier of autonomous mobile robotics is to enable agents to operate effectively in the wild in unseen unstructured environments. A significant challenge in unseen unstructured environments is that it may not be possible to predict all the dangers to the specific robot. Although recent work has used large foundation vision-language models (VLMs) to preemptively predict an exhaustive list of common-sense dangers, it remains difficult to capture possible interaction and embodiment-dependent adversities. We propose a continual learning framework for a mobile embodied agent to learn online from disturbances and attribute anomalous behaviours to causes through semantics, enabling better prediction and planning of the world in the future. Our framework, "Don't Fool Me Twice", first observes disturbances and describes their effects on the robot; this description is augmented with visual context to query a VLM to predict possible causes; the local disturbance is characterized using kernel regression, which allows for efficient, few-shot modeling of transient anomalies. We leverage semantic voxel-centric modeling to estimate epistemic uncertainty, enabling richer downstream recovery by treating interaction-driven disturbances as learnable spatial behaviors. We present four hypotheses and validate them in simulation and on hardware across embodiments and adversity modes.

机器人持续学习具身智能异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。