让智能体像老师一样动态调整难度,自动生成学习课程。
Heterogeneous Adversarial Play in Interactive Environments
- 用对抗机制让教师和学生角色互相进化,自动调节任务难度。
- 在多任务场景中表现媲美顶尖模型,提升人与智能体的学习效率。
- 适合需要自适应教学的开放场景,如机器人训练或个性化学习。
自对弈是自主技能习得的核心范式,智能体通过自我驱动的环境探索不断提升能力。传统自对弈框架依赖零和竞争中的对称性,但在存在固有不对称性的开放式学习场景中效果有限。人类教学系统正是非对称引导的典范,教育者根据学习者的发展轨迹设计适配挑战。核心难题在于如何在无人工预设任务层级的情况下,使人工智能系统自主构建合适的课程。本文提出异质对抗式学习(Heterogeneous Adversarial Play, HAP),将师生互动形式化为一个极小极大优化问题,其中生成任务的指导者与解决问题的学习者通过对抗动态共同演化。不同于现有自动课程学习方法采用静态课程或单向任务选择,HAP建立双向反馈机制,指导者依据学习者实时表现持续调整任务复杂度。在多个多任务学习领域上的实验验证表明,该框架性能达到当前最优水平,同时生成的课程显著提升了人工代理与人类受试者的学习成效。
原文摘要 · Abstract (English)
Self-play constitutes a fundamental paradigm for autonomous skill acquisition, whereby agents iteratively enhance their capabilities through self-directed environmental exploration. Conventional self-play frameworks exploit agent symmetry within zero-sum competitive settings, yet this approach proves inadequate for open-ended learning scenarios characterized by inherent asymmetry. Human pedagogical systems exemplify asymmetric instructional frameworks wherein educators systematically construct challenges calibrated to individual learners' developmental trajectories. The principal challenge resides in operationalizing these asymmetric, adaptive pedagogical mechanisms within artificial systems capable of autonomously synthesizing appropriate curricula without predetermined task hierarchies. Here we present Heterogeneous Adversarial Play (HAP), an adversarial Automatic Curriculum Learning framework that formalizes teacher-student interactions as a minimax optimization wherein task-generating instructor and problem-solving learner co-evolve through adversarial dynamics. In contrast to prevailing ACL methodologies that employ static curricula or unidirectional task selection mechanisms, HAP establishes a bidirectional feedback system wherein instructors continuously recalibrate task complexity in response to real-time learner performance metrics. Experimental validation across multi-task learning domains demonstrates that our framework achieves performance parity with SOTA baselines while generating curricula that enhance learning efficacy in both artificial agents and human subjects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。