融合多维度信息提升机器人抓取食物的适应性与成功率。
IMRL: Integrating Visual, Physical, Temporal, and Geometric Representations for Enhanced Food Acquisition
- 整合视觉、物理、时间与几何特征,构建更全面的食物表征
- 在真实机器人上实现35%成功率提升,支持零样本泛化
- 适合需要精准抓取复杂食物的助老助残机器人研发
机器人辅助进食对改善吞咽障碍者的生活质量具有重要意义。然而,在不同条件下获取多样食物并泛化到未见食物仍面临挑战。现有方法依赖视觉线索(如颜色、形状、纹理)提取的表面几何信息(如边界框和姿态),缺乏适应性和鲁棒性,尤其当食物物理属性相似但外观不同时表现不佳。本文采用模仿学习(IL)构建食物获取策略。现有方法使用ResNet-50等预训练图像编码器,其表征不够稳健,难以跨场景泛化。为此,我们提出新型方法IMRL(集成多维表征学习),融合视觉、物理、时间与几何信息,以增强IL在食物获取中的鲁棒性与泛化能力。该方法能识别食物类型与物理状态(固态、半固态、颗粒状、液态及混合物),建模动作的时间动态,并引入几何信息以确定最佳舀取点与判断碗体满度。IMRL使机器人可根据上下文自适应调整舀取策略,显著提升应对复杂食物场景的能力。真实机器人实验表明,本方法在多种食物与碗型配置下均表现出强鲁棒性与适应性,实现对未见场景的零样本泛化,相较最优基线成功率达35%提升。
原文摘要 · Abstract (English)
Robotic assistive feeding holds significant promise for improving the quality of life for individuals with eating disabilities. However, acquiring diverse food items under varying conditions and generalizing to unseen food presents unique challenges. Existing methods that rely on surface-level geometric information (e.g., bounding box and pose) derived from visual cues (e.g., color, shape, and texture) often lacks adaptability and robustness, especially when foods share similar physical properties but differ in visual appearance. We employ imitation learning (IL) to learn a policy for food acquisition. Existing methods employ IL or Reinforcement Learning (RL) to learn a policy based on off-the-shelf image encoders such as ResNet-50. However, such representations are not robust and struggle to generalize across diverse acquisition scenarios. To address these limitations, we propose a novel approach, IMRL (Integrated Multi-Dimensional Representation Learning), which integrates visual, physical, temporal, and geometric representations to enhance the robustness and generalizability of IL for food acquisition. Our approach captures food types and physical properties (e.g., solid, semi-solid, granular, liquid, and mixture), models temporal dynamics of acquisition actions, and introduces geometric information to determine optimal scooping points and assess bowl fullness. IMRL enables IL to adaptively adjust scooping strategies based on context, improving the robot's capability to handle diverse food acquisition scenarios. Experiments on a real robot demonstrate our approach's robustness and adaptability across various foods and bowl configurations, including zero-shot generalization to unseen settings. Our approach achieves improvement up to $35\%$ in success rate compared with the best-performing baseline. More details can be found on our website https://ruiiu.github.io/imrl.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。