让机器人像人一样通过互动理解环境,自主决策动作。
INTENTION: Inferring Tendencies of Humanoid Robot Motion Through Interactive Intuition and Grounded VLM
- 用视觉语言模型+交互记忆图实现环境理解与行为推理
- 无需重复指令即可在新场景中推断合理操作方式
- 适合需要灵活适应的现实场景机器人应用
传统机器人操控依赖精确物理模型和预设动作序列,在结构化环境中有效,但在真实场景中因建模误差难以泛化。人类则凭借直观的交互能力,基于隐含物理认知高效决策。本文提出INTENTION框架,通过结合视觉语言模型(VLM)的场景推理与交互驱动的记忆机制,使机器人具备学习性交互直觉。引入记忆图(Memory Graph)记录过往任务交互中的场景信息,体现类人对任务的理解与决策;设计直觉感知器(Intuitive Perceptor)从视觉场景中提取物理关系与可操作性。两者协同使机器人在无重复指令情况下,推断新场景下的合理交互行为。视频展示:https://robo-intention.github.io
原文摘要 · Abstract (English)
Traditional control and planning for robotic manipulation heavily rely on precise physical models and predefined action sequences. While effective in structured environments, such approaches often fail in real-world scenarios due to modeling inaccuracies and struggle to generalize to novel tasks. In contrast, humans intuitively interact with their surroundings, demonstrating remarkable adaptability, making efficient decisions through implicit physical understanding. In this work, we propose INTENTION, a novel framework enabling robots with learned interactive intuition and autonomous manipulation in diverse scenarios, by integrating Vision-Language Models (VLMs) based scene reasoning with interaction-driven memory. We introduce Memory Graph to record scenes from previous task interactions which embodies human-like understanding and decision-making about different tasks in real world. Meanwhile, we design an Intuitive Perceptor that extracts physical relations and affordances from visual scenes. Together, these components empower robots to infer appropriate interaction behaviors in new scenes without relying on repetitive instructions. Videos: https://robo-intention.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。