让机器人通过视觉触觉感知,动态判断夹、舀等动作对食物的适用性。
SAVOR: Skill Affordance Learning from Visuo-Haptic Perception for Robot-Assisted Bite Acquisition
- 结合工具功能与食物物理特性,构建可动态更新的动作适用性模型。
- 在20种单食材和10餐实际进餐中,成功率比现有方法提升13%。
- 适合需要自适应进食辅助的机器人系统研发者使用。
机器人辅助进食需可靠完成取食动作,但餐具与食物间复杂交互及食物属性随时间变化(如牛排冷却变硬)带来挑战。为此,我们提出SAVOR,一种从视觉-触觉感知中学习取食技能适配性的新方法。技能适配性由工具适配性(餐具能做什么)与食物适配性(食物允许什么)共同决定。工具适配性通过离线标定获得,涵盖多种餐具与食物组合的功能建模;食物适配性则基于软度、含水率、黏稠度等物理属性,先由视觉条件语言模型推断,再通过SAVOR-Net在交互中实时融合多模态感知动态优化。该方法在线融合离线与在线估计,实现技能适配性的实时预测,使机器人可为每类食物选择最优操作。在20种单食材及10次真实用餐场景测试中,较当前最佳类别基方法(如水果用叉子)提升13%取食成功率,验证了交互驱动技能适配性对泛化与高效取食的重要性。
原文摘要 · Abstract (English)
Robot-assisted feeding requires reliable bite acquisition, a challenging task due to the complex interactions between utensils and food with diverse physical properties. These interactions are further complicated by the temporal variability of food properties-for example, steak becomes firm as it cools even during a meal. To address this, we propose SAVOR, a novel approach for learning skill affordances for bite acquisition-how suitable a manipulation skill (e.g., skewering, scooping) is for a given utensil-food interaction. In our formulation, skill affordances arise from the combination of tool affordances (what a utensil can do) and food affordances (what the food allows). Tool affordances are learned offline through calibration, where different utensils interact with a variety of foods to model their functional capabilities. Food affordances are characterized by physical properties such as softness, moisture, and viscosity, initially inferred through commonsense reasoning using a visually-conditioned language model and then dynamically refined through online multi-modal visuo-haptic perception using SAVOR-Net during interaction. Our method integrates these offline and online estimates to predict skill affordances in real time, enabling the robot to select the most appropriate skill for each food item. Evaluated on 20 single-item foods and 10 in-the-wild meals, our approach improves bite acquisition success rate by 13% over state-of-the-art (SOTA) category-based methods (e.g. use skewer for fruits). These results highlight the importance of modeling interaction-driven skill affordances for generalizable and effective robot-assisted bite acquisition. Website: https://emprise.cs.cornell.edu/savor/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。