arXiv:2607.24207cs.RO2026-07

让机器人更懂地面操作空间,提升抓取成功率。

FloAff-Kitchen: Bridging Navigation and Manipulation via Canonical and Progressive Floor Affordance Learning

论文配图:FloAff-Kitchen: Bridging Navigation and Manipulation via Canonical and Progressive Floor Affordance Learning
图 1 · 摘自论文原文
  • 用统一框架学习可迁移的地面操作空间表示
  • 在多个场景中比现有方法提升15%以上成功率
  • 适合需要精准操作的家务机器人研究者

移动操作需识别能最大化下游操作成功率的地面操作空间(FloAff),而非仅保证导航可行性。现有方法因无关空间上下文和物体任意朝向导致表征模糊,并混淆共享与任务特异性知识。为此,我们提出统一框架,从第一人称多模态感知中进行范式化表征与渐进式先验学习。引入范式地面操作表征(CFAR),通过保留与操作相关的局部结构并消除与机器人基座位置无关的干扰变化,学习标准化交互几何。进一步提出渐进式地面操作学习(PFAL),从基础操作任务中学习可迁移的FloAff先验,并逐步适配多样化的下游操作技能。为系统评估,我们建立首个跨场景、多视角的FloAff-Kitchen基准,覆盖多种操作技能、场景布局、家具风格和视角。三个设置下的大量实验表明,本方法持续优于强基线;消融实验证实各组件有效性。

原文摘要 · Abstract (English)

Mobile manipulation requires robots to identify Floor Affordance (FloAff) that maximizes downstream manipulation success rather than merely ensuring navigation feasibility. FloAff prediction is a target-conditioned local spatial reasoning problem, yet existing methods suffer from representation ambiguity caused by irrelevant spatial context and arbitrary object orientations, while entangling shared and task-specific knowledge across heterogeneous manipulation skills. To address these challenges, we propose a unified framework for FloAff prediction from egocentric multimodal perception, consisting of canonical representation learning and progressive affordance prior learning. Specifically, we introduce a Canonical Floor Affordance Representation (CFAR), which learns canonical interaction geometry by preserving affordance-relevant local structure while eliminating nuisance spatial variations unrelated to robot base placement. We further propose Progressive Floor Affordance Learning (PFAL), which learns transferable FloAff priors from a foundation manipulation task and progressively adapts them to heterogeneous downstream manipulation skills. To facilitate systematic evaluation, we establish the first cross-scene, multi-view FloAff-Kitchen benchmark covering diverse manipulation skills, scene layouts, furniture styles, and viewpoints. Extensive experiments on three benchmark settings demonstrate that our method consistently outperforms strong baselines, while ablation studies validate the contribution of each proposed component. Project page: https://csu-hero-lab.github.io/FloAff-Kitchen_Web/

移动操作地面操作多模态感知机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。