让机器人根据指令精准判断物品可操作区域,提升智能操作能力
Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model
- 用指令驱动的视角预测物品可操作区域,突破传统固定预设
- 构建1.5万组指令-物体-可操作性三元组数据集,覆盖第一人称视角
- 利用大模型自验证推理机制,实现高精度、可解释的指令导向预测
在物体操作任务中,可操作性对智能机器人至关重要。本文指出,可操作性应依赖于任务或指令,而这一关键特性被以往多数工作忽略:同一物体在不同指令下可能对应不同的操作区域与方向。基于此,我们构建了一个包含1.5万组物体-指令-可操作性三元组的新数据集,所有场景均来自第一人称视角,模拟类人机器人视角。此外,我们探索如何利用大模型(LMMs)作为可操作性预测器,提出一种“搜索对抗验证”流程:让大模型逐步生成可操作性预测,每步输出由自身验证,模拟推理过程。实验表明,该方法不仅实现了全新的指令导向可操作性预测能力,且在多个指标上表现优异。
原文摘要 · Abstract (English)
Affordance is crucial for intelligent robots in the context of object manipulation. In this paper, we argue that affordance should be task-/instruction-dependent, which is overlooked by many previous works. That is, different instructions can lead to different manipulation regions and directions even for the same object. According to this observation, we present a new dataset comprising fifteen thousand object-instruction-affordance triplets. All scenes in the dataset are from an egocentric viewpoint, designed to approximate the perspective of a human-like robot. Furthermore, we investigate how to enable large multimodal models (LMMs) to serve as affordance predictors by implementing a ``search against verifiers'' pipeline. An LMM is asked to progressively predict affordances, with the output at each step being verified by itself during the iterative process, imitating a reasoning process. Experiments show that our method not only unlocks new instruction-oriented affordance prediction capabilities, but also achieves outstanding performance broadly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。