arXiv:2505.00439cs.LGcs.AI2025-05被引 2

动态验证集提升图神经网络策略的跨规模泛化能力

Per-Domain Generalizing Policies: On Validation Instances and Scaling Behavior

  • 动态生成随训练进程增大的验证实例,避免固定验证集偏差
  • 在9个领域中均显著改善图神经网络策略的规模扩展性能
  • 适合研究模型泛化性与大规模强化学习的学者参考

近期研究表明,可在特定领域内学习出具备泛化能力的动作策略。其核心目标是实现从小规模训练实例到大规模测试实例的可扩展性;而使用比训练实例更大的验证实例,是达成该目标的关键。以往工作采用固定验证集,本文提出一种动态生成验证集的方法,在信息充足且可行的前提下实时增大实例规模。同时引入改进的评估方法,系统生成测试实例以确保每个实例规模下性能覆盖率具有给定置信度。实验表明,在所用的9个领域中,动态验证均显著提升了图神经网络策略的可扩展性。

原文摘要 · Abstract (English)

Recent work has shown that successful per-domain generalizing action policies can be learned. Scaling behavior, from small training instances to large test instances, is the key objective; and the use of validation instances larger than training instances is one key to achieve it. Prior work has used fixed validation sets. Here, we introduce a method generating the validation set dynamically, on the fly, increasing instance size so long as informative and feasible.We also introduce refined methodology for evaluating scaling behavior, generating test instances systematically to guarantee a given confidence in coverage performance for each instance size. In experiments, dynamic validation improves scaling behavior of GNN policies in all 9 domains used.

图神经网络策略泛化可扩展性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。