arXiv:2512.24145cs.LGstat.ML2025-12

配对种子可降低评估方差,提升对比实验的统计效力。

When Does Pairing Seeds Reduce Variance? Evidence from a Multi-Agent Economic Simulation

  • 用相同随机种子配对评估,使不同系统在相同随机条件下运行。
  • 在固定预算下,配对评估能发现独立评估无法察觉的系统差异。
  • 适用于需要高精度对比的强化学习或多智能体仿真场景。

机器学习系统看似随机,实则由确定性伪随机数生成器驱动,重复运行时产生完全相同的样本。当前评估通常将不同配置的运行视为独立,未利用共享随机源的优势。本文分析了在共享随机种子下的比较评估统计结构:当不同系统使用相同种子时,其随机结果被配对,只要结果在种子层面正相关,即可严格减少方差。我们通过一个扩展的学习型多智能体经济模拟器验证了该效应:配对评估揭示了聚合与分布层面的系统性差异,而独立评估在固定预算下仍无法得出显著结论。

原文摘要 · Abstract (English)

Machine learning systems appear stochastic but are deterministically random, as seeded pseudorandom number generators produce identical realisations across repeated executions. Standard evaluation practice typically treats runs across alternatives as independent and does not exploit shared sources of randomness. This paper analyses the statistical structure of comparative evaluation under shared random seeds. Under this design, competing systems are evaluated using identical seeds, inducing matched stochastic realisations and yielding strict variance reduction whenever outcomes are positively correlated at the seed level. We demonstrate these effects using an extended learning-based multi-agent economic simulator, where paired evaluation exposes systematic differences in aggregate and distributional outcomes that remain statistically inconclusive under independent evaluation at fixed budgets.

多智能体评估方法方差控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。