arXiv:2608.04042cs.RO2026-08

用基础模型搭建厨房机器人感知系统,实现无重训操作。

Kitchen Robotic Manipulation utilizing Foundation Models

论文配图:Kitchen Robotic Manipulation utilizing Foundation Models
图 1 · 摘自论文原文
  • 模块化设计融合多模型,支持灵活替换与优化
  • 在20个场景下达89.12%的ADI,应对杂乱遮挡
  • 真实机器人验证,可直接部署不需环境微调

将机器人部署于日常人类环境需要具备鲁棒且适应性强的感知系统。本文提出一种面向家庭操作任务的模块化感知流程,聚焦厨房餐具处理。该流程集成开放词汇物体检测、多视角分割、实例感知3D重建以及2D-3D特征融合策略,用于6D位姿估计与抓取规划。模块化结构支持多种视觉与几何基础模型的系统性替换,通过在自建厨房数据集上的大量评估,确定最优配置(LLMDet + SAMv2 + DINOv2 + GeoTransformer)。该配置在20场景厨房基准测试中,于杂乱与遮挡条件下取得89.12%的ADI。真实世界演示表明,该配置可在物理机器人上直接部署,无需环境特定微调,成功完成从水槽到洗碗机转移及杯子堆叠等任务,验证了系统的适应性与可扩展性。代码与补充材料见 https://raivlab.github.io/FM_kitchen。

原文摘要 · Abstract (English)

Deploying robots in everyday human environments requires perception systems that are both robust and adaptable to diverse, dynamic conditions. In this work, we present a modular perception pipeline for household manipulation tasks, with a focus on dishware handling in kitchen environments. The pipeline integrates open-vocabulary object detection, multi-view segmentation, instance-aware 3D reconstruction, and a 2D-3D feature fusion strategy for 6D pose estimation and grasp planning. Its modular design enables systematic substitution of multiple visual and geometric foundation models, allowing us to identify the best-performing configuration through extensive evaluation on a custom kitchen dataset. The best-performing configuration (LLMDet + SAMv2 + DINOv2 + GeoTransformer) achieves an ADI of 89.12\% on the 20-scene kitchen benchmark with cluttered and occluded conditions. Furthermore, real-world demonstrations confirm that the best configuration can be deployed on physical robots without environment-specific retraining, successfully executing tasks such as sink-to-dishwasher transfer and cup stacking. It validates the adaptability and scalability of the pipeline and highlights its potential as a practical framework for household robotic systems. Our code and supplementary materials are available at https://raivlab.github.io/FM_kitchen .

机器人感知基础模型厨房机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。