让机器人通过主动提问高效学习物品归属关系。
Toward Ownership Understanding of Objects: Active Question Generation with Large Language Model and Probabilistic Generative Model
- 结合大模型常识与概率生成模型,智能选择提问对象。
- 用更少问题实现更高归属识别准确率,实验验证有效。
- 适合需要理解物品归属的居家/办公机器人场景。
在家庭和办公室环境中,机器人需理解物品归属才能正确执行如“把我的杯子拿来”等指令。然而,仅凭视觉特征无法可靠推断归属。为此,我们提出主动归属学习(ActOwL)框架,使机器人能主动生成并询问归属相关问题。该框架采用概率生成模型,选择信息增益最大的问题以高效获取归属知识。同时,借助大语言模型(LLM)的常识知识,预先将物品分类为共享或私有,仅对私有物品进行提问。在模拟家庭环境和真实实验室中进行的实验表明,与基线方法相比,ActOwL在更少提问次数下实现了显著更高的归属聚类准确率。结果证明,结合主动推理与大模型引导的常识推理,可有效提升机器人获取归属知识的能力,支持更实用、更符合社会规范的任务执行。
原文摘要 · Abstract (English)
Robots operating in domestic and office environments must understand object ownership to correctly execute instructions such as ``Bring me my cup.'' However, ownership cannot be reliably inferred from visual features alone. To address this gap, we propose Active Ownership Learning (ActOwL), a framework that enables robots to actively generate and ask ownership-related questions to users. ActOwL employs a probabilistic generative model to select questions that maximize information gain, thereby acquiring ownership knowledge efficiently to improve learning efficiency. Additionally, by leveraging commonsense knowledge from Large Language Models (LLM), objects are pre-classified as either shared or owned, and only owned objects are targeted for questioning. Through experiments in a simulated home environment and a real-world laboratory setting, ActOwL achieved significantly higher ownership clustering accuracy with fewer questions than baseline methods. These findings demonstrate the effectiveness of combining active inference with LLM-guided commonsense reasoning, advancing the capability of robots to acquire ownership knowledge for practical and socially appropriate task execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。