arXiv:2604.25897cs.ROcs.LG2026-04中稿 · publication at IRO…被引 2

用可微分的高斯混合模型提升抓取在不确定下的鲁棒性。

Variational Neural Belief Parameterizations for Robust Dexterous Grasping under Multimodal Uncertainty

论文配图:Variational Neural Belief Parameterizations for Robust Dexterous Grasping under Multimodal Uncertainty
图 1 · 摘自论文原文
  • 用变分推断构建可微的接触参数信念,替代传统粒子滤波。
  • 在模拟中抓取成功率提升,规划时间缩短近10倍。
  • 适合需要快速、可靠抓取的机器人控制场景。

接触变化、感知不确定性与外部干扰使抓取执行具有随机性。传统期望质量目标忽略尾部风险,常选在不利接触条件下失败的抓取。风险敏感的部分可观马尔可夫决策过程(POMDP)可缓解此问题,但多数采用粒子滤波信念,存在扩展性差、难以梯度优化及条件风险价值(CVaR)估计方差高的缺陷。本文将抓取获取建模为对潜在接触参数与物体位姿的变分推断,用可微的高斯混合分布表示信念。通过Gumbel-Softmax选择组件和位置-尺度重参数化,将采样表达为信念参数的光滑函数,从而实现路径梯度,直接优化尾部鲁棒性的可微分CVaR代理。在仿真中,该方法在接触参数不确定性和外力扰动下显著提升抓取成功率,规划时间相比粒子滤波模型预测控制降低约一个数量级。在串列机械臂配多指手的实际系统上,验证了在物体位姿不确定性下的抓取-举升成功,相较高斯基线方法终止步数更少、实际耗时更低,且触觉抓取质量代理更高。所学信念风险校准更准确,平均绝对校准误差低于0.14,优于交叉熵方法的0.58。代码、仿真资源与243个力闭合抓取数据集已开源。

原文摘要 · Abstract (English)

Contact variability, sensing uncertainty, and external disturbances make grasp execution stochastic. Expected-quality objectives ignore tail outcomes and often select grasps that fail under adverse contact realizations. Risk-sensitive POMDPs address this failure mode, but many use particle-filter beliefs that scale poorly, obstruct gradient-based optimization, and estimate Conditional Value-at-Risk (CVaR) with high-variance approximations. We instead formulate grasp acquisition as variational inference over latent contact parameters and object pose, representing the belief with a differentiable Gaussian mixture. We use Gumbel-Softmax component selection and location-scale reparameterization to express samples as smooth functions of the belief parameters, enabling pathwise gradients through a differentiable CVaR surrogate for direct optimization of tail robustness. In simulation, our variational neural belief improves robust grasp success under contact-parameter uncertainty and exogenous force perturbations while reducing planning time by roughly an order of magnitude relative to particle-filter model-predictive control. On a serial-chain robot arm with a multifingered hand, we validate grasp-and-lift success under object-pose uncertainty against a Gaussian baseline. Both methods succeed on the tested perturbations, but our controller terminates in fewer steps and less wall-clock time while achieving a higher tactile grasp-quality proxy. Our learned belief also calibrates risk more accurately, keeping mean absolute calibration error below 0.14 across tested simulation regimes, compared with 0.58 for a Cross-Entropy Method planner. We provide code, simulation assets, and a dataset of 243 force-closed grasps at the following link: www.github.com/coenwerem/vnb-grasp.

抓取强化学习不确定性可微分推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。