现有复杂度度量方法在机器人任务中存在误导,需改进。
Issues with Measuring Task Complexity via Random Policies in Robotic Tasks
- 用随机策略表现评估任务复杂度,依赖随机权重猜测(RWG)
- PIC和POIC指标显示双连杆比单连杆简单,与实际相反
- 稀疏奖励任务被误判为比密集奖励简单,不契合实际经验
强化学习(RL)在机器人和自然语言处理等领域取得重大进展。衡量任务复杂度是构建有意义基准和设计有效课程的关键挑战。尽管表格化设置已有成熟度量方法,非表格领域仍相对匮乏。现有方法包括:(i) 通过随机权重猜测(RWG)分析随机策略的性能;(ii) 基于RWG的信息论指标,如策略信息容量(PIC)和最优策略信息容量(POIC)。本文在逐步复杂的机器人操作任务中评估这些方法,采用已知相对复杂度的设定,并对比密集与稀疏奖励形式。实验结果表明,复杂度测量仍具复杂性:相同奖励形式下,PIC认为双连杆机械臂比单连杆简单,这与机器人控制及实证RL认知相悖;同一任务中,POIC估计稀疏奖励任务比密集奖励更简单。因此,我们揭示了PIC和POIC与常规理解及实证结果矛盾,强调需超越基于RWG的度量,发展更可靠的非表格强化学习复杂度评估方法,本研究框架可作为起点。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has enabled major advances in fields such as robotics and natural language processing. A key challenge in RL is measuring task complexity, which is essential for creating meaningful benchmarks and designing effective curricula. While there are numerous well-established metrics for assessing task complexity in tabular settings, relatively few exist in non-tabular domains. These include (i) Statistical analysis of the performance of random policies via Random Weight Guessing (RWG), and (ii) information-theoretic metrics Policy Information Capacity (PIC) and Policy-Optimal Information Capacity (POIC), which are reliant on RWG. In this paper, we evaluate these methods using progressively difficult robotic manipulation setups, with known relative complexity, with both dense and sparse reward formulations. Our empirical results reveal that measuring complexity is still nuanced. Specifically, under the same reward formulation, PIC suggests that a two-link robotic arm setup is easier than a single-link setup - which contradicts the robotic control and empirical RL perspective whereby the two-link setup is inherently more complex. Likewise, for the same setup, POIC estimates that tasks with sparse rewards are easier than those with dense rewards. Thus, we show that both PIC and POIC contradict typical understanding and empirical results from RL. These findings highlight the need to move beyond RWG-based metrics towards better metrics that can more reliably capture task complexity in non-tabular RL with our task framework as a starting point.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。