让机器人通过属性组合精准找回物体,提升长期记忆能力
FRAME: Factored Retrieval via Attribute Readouts for Object-Centric Scene Memory

- 将语言指令转为属性权重,用读出机制提取物体特征证据
- 在多场景测试中超越现有基线,推理速度更快
- 适合需要长期记忆与多属性检索的具身智能系统
语言引导的机器人需要持久化场景记忆以执行指令、重访物体并解析时间跨度内的指代。日常物体指代常依赖多个持久属性(如类别、材质、大小、表面外观)。本文提出属性组合检索任务,即通过自然语言查询固定对象中心的场景记忆,定位满足所有指定属性的物体。为此,我们设计了基于固定场景记忆和属性定义目标的受控评估协议,排除感知与标注歧义干扰。提出FRAME方法:将语言转化为查询相关属性权重,通过学习到的读出机制从物体嵌入中估计各属性证据,并按查询聚合排名。在未见场景与物体资产上,FRAME显著优于代表性基线,同时将后分解物体评分降至轻量级矩阵-向量运算。结果表明,属性组合检索是语言引导机器人的重要补充能力,持久属性可作为可组合证据,实现准确高效多属性检索。
原文摘要 · Abstract (English)
Language-guided robots need persistent scene memories to follow instructions, revisit objects, and resolve references to objects encountered over time. While much of language-guided scene-memory retrieval has emphasized spatial or relational references, many everyday object references specify objects by multiple persistent attributes, such as category, material, size, or surface appearance. We formalize this problem as attribute-compositional retrieval, where a fixed object-centric scene memory is queried with natural language to retrieve the object satisfying the requested attributes. To investigate this capability directly, we introduce a controlled evaluation protocol with fixed scene memories and attribute-defined targets, separating retrieval from perception and annotation ambiguities. We then propose FRAME, which turns language into query-relevant attribute weights, uses learned readouts to estimate per-attribute evidence from object embeddings, and ranks objects by aggregating this evidence according to the query. Across held-out scenes and object assets, FRAME outperforms representative scene-memory retrieval baselines while reducing post-decomposition object scoring to lightweight matrix-vector computation. These results position attribute-compositional retrieval as a complementary scene-memory capability for language-guided robots, showing that persistent object attributes can be exposed as composable evidence for accurate and efficient multi-attribute retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。