提出多属性参考理解框架,提升机器人对自然指令的精准响应
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
- 融合状态、意图和手势的多属性参考定位
- 在新数据集上表现优于单属性方法
- 适合人机交互与智能机器人场景
指代表达理解(REC)旨在根据自然语言描述实现目标定位。然而现有方法受限于物体类别描述和单一属性意图描述,难以应用于真实场景。在人机交互中,用户常通过个体状态、意图及引导手势表达需求,而非详细物体描述。为此,我们提出多属性参考理解(Multi-ref EC)框架,整合状态、意图与具身手势以定位目标。构建了包含状态、意图表达与具身参照的SIGAR数据集。在多种基线模型上的实验表明,合理排序的多属性参考能显著提升定位性能,证明单属性参考不足以应对自然人机交互场景。研究结果强调多属性参考在视觉-语言理解中的重要性。
原文摘要 · Abstract (English)
Referring expression comprehension (REC) aims at achieving object localization based on natural language descriptions. However, existing REC approaches are constrained by object category descriptions and single-attribute intention descriptions, hindering their application in real-world scenarios. In natural human-robot interactions, users often express their desires through individual states and intentions, accompanied by guiding gestures, rather than detailed object descriptions. To address this challenge, we propose Multi-ref EC, a novel task framework that integrates state descriptions, derived intentions, and embodied gestures to locate target objects. We introduce the State-Intention-Gesture Attributes Reference (SIGAR) dataset, which combines state and intention expressions with embodied references. Through extensive experiments with various baseline models on SIGAR, we demonstrate that properly ordered multi-attribute references contribute to improved localization performance, revealing that single-attribute reference is insufficient for natural human-robot interaction scenarios. Our findings underscore the importance of multi-attribute reference expressions in advancing visual-language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。