改进机器人领域中参数推断的分布支持,提升不确定性建模可靠性。
Heuristic Adaptation of Potentially Misspecified Domain Support for Likelihood-Free Inference in Stochastic Dynamical Systems
- 提出三种启发式方法动态调整先验支持以应对分布误设问题。
- 在可变形物体操控任务中实现更精细的长度与刚度分类精度。
- 适用于需要高鲁棒性的仿真训练策略学习场景。
在机器人领域,无似然推断(LFI)可为部署条件参数集提供适配的领域分布。传统LFI假设采样支持固定不变,随迭代优化从初始先验逐步逼近后验分布。然而,若支持设定错误,可能导致次优但看似确定的后验结果。为此,本文提出三种启发式LFI变体:EDGE、MODE和CENTRE,各自以不同方式解释后验众数随推断步骤的变化,并在每一步中协同调整支持范围。首先揭示了支持误设的问题,在随机动力系统基准上评估所提启发式方法;随后在可变形线性物体(DLO)操控任务中测试其对参数推断与策略学习的影响。结果表明,该方法能实现更精细的长度与刚度分类;当使用生成的后验作为仿真策略学习的领域分布时,显著提升以对象为中心的智能体性能鲁棒性。
原文摘要 · Abstract (English)
In robotics, likelihood-free inference (LFI) can provide the domain distribution that adapts a learnt agent in a parametric set of deployment conditions. LFI assumes an arbitrary support for sampling, which remains constant as the initial generic prior is iteratively refined to more descriptive posteriors. However, a potentially misspecified support can lead to suboptimal, yet falsely certain, posteriors. To address this issue, we propose three heuristic LFI variants: EDGE, MODE, and CENTRE. Each interprets the posterior mode shift over inference steps in its own way and, when integrated into an LFI step, adapts the support alongside posterior inference. We first expose the support misspecification issue and evaluate our heuristics using stochastic dynamical benchmarks. We then evaluate the impact of heuristic support adaptation on parameter inference and policy learning for a dynamic deformable linear object (DLO) manipulation task. Inference results in a finer length and stiffness classification for a parametric set of DLOs. When the resulting posteriors are used as domain distributions for sim-based policy learning, they lead to more robust object-centric agent performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。