提出自适应信息目标QOED,让机器人高效探索关键参数。
Learning What Matters: Adaptive Information-Theoretic Objectives for Robot Exploration

- 基于费雪信息矩阵分析,筛选可识别参数方向。
- 抑制非关键参数干扰,提升探索效率35.23%。
- 适合需要高效数据采集的机器人学习任务。
设计可学习的信息论目标以指导机器人探索仍具挑战性。这类目标旨在引导探索以减少模型参数的不确定性,但往往难以确定所收集数据实际能揭示的信息。尽管强化学习(RL)可优化给定目标,但在高维机器人系统中构建反映参数可学习性的目标十分困难。许多参数方向可观测性弱或不可识别,即使选择了可识别方向,忽略的方向仍可能影响探索并扭曲信息度量。为此,我们提出拟最优实验设计(QOED),一种基于最优实验设计的自适应信息目标。QOED(i)对费雪信息矩阵进行特征空间分析,识别可观测子空间并选择可识别的参数方向;(ii)修改探索目标,强调这些方向的同时抑制非关键参数的干扰。在有限干扰和关键与非关键方向耦合较弱的前提下,QOED提供了对理想信息目标的常数因子近似,该目标可探索所有参数。我们在模拟和真实世界的导航与操作任务中评估了QOED,其中可识别方向选择和噪声抑制分别带来35.23%和21.98%的性能提升。当作为基于模型的策略优化中的探索目标集成时,QOED进一步优于现有强化学习基线。
原文摘要 · Abstract (English)
Designing learnable information-theoretic objectives for robot exploration remains challenging. Such objectives aim to guide exploration toward data that reduces uncertainty in model parameters, yet it is often unclear what information the collected data can actually reveal. Although reinforcement learning (RL) can optimize a given objective, constructing objectives that reflect parametric learnability is difficult in high-dimensional robotic systems. Many parameter directions are weakly observable or unidentifiable, and even when identifiable directions are selected, omitted directions can still influence exploration and distort information measures. To address this challenge, we propose Quasi-Optimal Experimental Design (Q{\footnotesize OED}), an adaptive information objective grounded in optimal experimental design. Q{\footnotesize OED} (i) performs eigenspace analysis of the Fisher information matrix to identify an observable subspace and select identifiable parameter directions, and (ii) modifies the exploration objective to emphasize these directions while suppressing nuisance effects from non-critical parameters. Under bounded nuisance influence and limited coupling between critical and nuisance directions, Q{\footnotesize OED} provides a constant-factor approximation to the ideal information objective that explores all parameters. We evaluate Q{\footnotesize OED} on simulated and real-world navigation and manipulation tasks, where identifiable-direction selection and nuisance suppression yield performance improvements of \SI{35.23}{\percent} and \SI{21.98}{\percent}, respectively. When integrated as an exploration objective in model-based policy optimization, Q{\footnotesize OED} further improves policy performance over established RL baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。