让机器人更懂模糊指令,精准定位现实场景中的物体。
SaaF: Scene-Specific Ambiguity-Aware 3D Language Fields towards Interactive Real-World Object Retrieval

- 用多视角特征拉近法增强物体间视觉区分度
- 联合训练多类描述词,识别并响应模糊查询
- 适合需要自然语言交互的智能服务机器人
我们提出场景特定的模糊感知3D语言场(SaaF),一种基于高斯泼溅的3D语言场,用于在真实场景中实现交互式物体检索。当前3D语言场方法通过渲染像素与自编码器压缩的CLIP特征建立关联,存在两个缺陷:(1) 特征压缩导致相似物体区分能力下降;(2) 对模糊查询处理差,常引发不稳定或错误检索。SaaF引入度量学习构建统一特征空间,兼具实例可区分性与模糊感知能力:(i) 通过度量学习将同一物体多视角图像特征在特征空间中拉近,提升实例级视觉区分度;(ii) 联合训练由所提方法生成的每段追踪物体图像序列的多个文本标签(含模糊描述),学习目标场景中模糊与具体特征间的语义关系。该特征空间支持细粒度视觉理解,并可估计查询模糊性,必要时交互请求澄清。实验表明,SaaF在开放词汇设置下不仅提升检索准确率,还稳健检测与处理用户文本查询中的模糊性。
原文摘要 · Abstract (English)
We propose Scene-specific Ambiguity-aware 3D Language Fields (SaaF), a novel Gaussian Splatting-based 3D language field designed for interactive object retrieval in a given real-world scene. Interactive object retrieval using natural language is a crucial capability for service robots operating in complex real-world environments. While recent 3D language field methods for object retrieval establish associations between rendered pixels and autoencoder-compressed CLIP features, they suffer from two limitations: (1) reduced discriminability among similar objects due to feature compression, and (2) poor handling of ambiguous queries, often resulting in unstable or incorrect retrieval. To address these limitations, SaaF introduces a metric learning strategy to construct a unified feature space that is both instance-discriminative and ambiguity-aware. (i) To enhance instance-level visual discrimination, SaaF employs metric learning that pulls image features from multiple viewpoints of the same object closer together in the feature space. (ii) To establish ambiguity awareness, the model jointly trains on multiple text labels generated by the proposed method from each tracked object image sequence, including ambiguous descriptions, to learn the semantic relationships between ambiguous and specific features in a target scene. This feature space enables fine-grained visual understanding while allowing the system to estimate query ambiguity and interactively request clarification when needed. Experimental results demonstrate that SaaF not only improves retrieval accuracy over previous methods but also robustly detects and handles ambiguity in the user text queries under open-vocabulary settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。