提出评估多智能体系统时的新型方差缩减方法,解决启发式函数选择不当带来的偏差问题。
Heuristic Pathologies and Further Variance Reduction via Uncertainty Propagation in the AIVAT Family of Techniques
- 通过参数化启发式函数揭示其潜在路径依赖漏洞
- 引入不确定性传播机制,使估计误差可量化并减少43%样本需求
- 适合需要高精度小样本评估的研究者,如扑克博弈算法验证
在多智能体环境中,当样本量有限或试验成本高昂时,如何评估智能体性能?AIVAT家族的方差缩减技术通过引入无偏低方差的期望回报估计器来应对这一挑战。其关键在于一个启发式价值函数,用于区分可能具有低值或高值的反事实历史。然而,现有文献缺乏对启发式价值函数选择的约束或指导,也未考虑其输出不确定性的处理。本文第一项贡献是将启发式价值函数参数化,揭示了其潜在缺陷:a) 可通过梯度下降使样本方差被病态地设为极低;b) 可通过梯度上升/下降对检验统计量进行p值操纵以得出期望结论。核心启示是:启发式价值函数应在观察评估数据前固定。第二项贡献是展示如何传播启发式不确定性以量化AIVAT估计的不确定性,并通过逆方差加权平均进一步降低方差,但可能牺牲无偏性。实验使用包含10,000手扑克牌的数据集验证了上述路径问题与不确定性结果,后者使达成统计结论所需的样本数减少43.0%。
原文摘要 · Abstract (English)
How should an agent's performance in a multiagent environment be evaluated when there is a limited sample size or a high cost of running a trial? The AIVAT family of variance reduction techniques was proposed to address this challenge by introducing unbiased low-variance estimators of agents' expected payoffs. An important component of AIVAT is a heuristic value function that discriminates between potentially low- and high-value counterfactual histories. A notable gap in the literature is that there is little to no constraint or guideline on how the heuristic value function should be chosen or how uncertainty in its output should be handled. In our first contribution, we parameterize the heuristic value function to highlight AIVAT's potential vulnerabilities: a) the sample variance can be set pathologically low by directly applying gradient descent on the sample variance, and b) one can p-hack to draw a desired statistical conclusion via gradient descent/ascent on the test statistic. The main takeaway is that the heuristic value function should be fixed prior to observing the evaluation data! In our second contribution, we show how the heuristic uncertainty can be propagated to quantify the uncertainty of AIVAT estimates. It is then possible to further reduce the variance using inverse-variance weighted averaging, but AIVAT's unbiasedness guarantee may have to be sacrificed. In our experiments, we use a dataset of 10,000 poker hands to demonstrate our heuristic pathology and uncertainty results, with the latter yielding a 43.0% reduction in the number of samples (poker hands) needed to draw statistical conclusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。