用语义结构与性能优化提升机器人精细操作能力
ReSemAct: Advancing Fine-Grained Robotic Manipulation via Semantic Structuring and Affordance Refinement
- 融合多模态大模型与视觉基础模型,构建语义化操作目标
- 零样本条件下在家庭与实验室环境完成多样化任务
- 适合需要精准语义理解的机器人操作研究者
精细机器人操作需将自然语言准确映射到具体可操作属性。现有基于基础模型的方法常将丰富语义简化为粗糙属性,导致隐含信息丢失。为此,我们提出ReSemAct,一种统一操作框架,引入语义结构化与属性精炼(SSAR)机制,依托多模态大模型(MLLMs)与视觉基础模型(VFMs)的协同推理。语义结构化模块从自然语言与RGB观测中提取统一语义属性描述,整合属性区域、隐式功能意图与粗粒度锚点,形成结构化表示以支持后续精炼。在此基础上,属性精炼策略设计两种互补流程,分别优化几何形状与空间位置,生成精细操作目标。这些目标被编码为实时关节空间优化目标,实现在动态环境中的反应式、鲁棒操作。在语义丰富的家庭环境与稀疏化学实验室环境中进行大量仿真与真实世界实验。结果表明,ReSemAct在零样本条件下完成多样任务,验证了基于基础模型的SSAR机制在精细操作中的鲁棒性。代码与视频见https://github.com/scy-v/ReSemAct和https://resemact.github.io。
原文摘要 · Abstract (English)
Fine-grained robotic manipulation requires grounding natural language into appropriate affordance targets. However, most existing methods driven by foundation models often compress rich semantics into oversimplified affordances, preventing exploitation of implicit semantic information. To address these challenges, we present ReSemAct, a novel unified manipulation framework that introduces Semantic Structuring and Affordance Refinement (SSAR), powered by the automated synergistic reasoning between Multimodal Large Language Models (MLLMs) and Vision Foundation Models (VFMs). Specifically, the Semantic Structuring module derives a unified semantic affordance description from natural language and RGB observations, organizing affordance regions, implicit functional intent, and coarse affordance anchors into a structured representation for downstream refinement. Building upon this specification, the Affordance Refinement strategy instantiates two complementary flows that separately specialize geometry and position, yielding fine-grained affordance targets. These refined targets are then encoded as real-time joint-space optimization objectives, enabling reactive and robust manipulation in dynamic environments. Extensive simulation and real-world experiments are conducted in semantically rich household and sparse chemical lab environments. The results demonstrate that ReSemAct performs diverse tasks under zero-shot conditions, showcasing the robustness of SSAR with foundation models in fine-grained manipulation. Code and videos at https://github.com/scy-v/ReSemAct and https://resemact.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。