让大模型实时根据优化反馈动态设计训练任务,提升多任务策略搜索效果。
Interactive LLM-assisted Curriculum Learning for Multi-Task Evolutionary Policy Search
- 大模型基于进化优化的实时反馈,动态生成训练任务
- 结合进度图与行为可视化反馈时,性能媲美人工设计课程
- 适用于具身智能与进化机器人等需要复杂训练策略的场景
多任务策略搜索面临挑战,因策略需在训练之外泛化。课程学习通过逐步增加难度已被证明有效,但设计优质课程耗时且依赖领域知识。基于大模型的课程生成虽为新方向,但以往仅限静态离线模式,未能利用优化器的实时反馈。本文提出一种交互式大模型辅助课程生成框架,使大模型根据进化优化过程的实时反馈自适应设计训练案例。研究不同反馈模态(仅数值指标、结合图表与行为可视化)对课程生成质量的影响。以遗传编程为优化器,在2D机器人导航任务中评估该方法,对比静态大模型生成课程与人工设计基线。结果表明,交互式课程生成优于静态方法,多模态反馈(含进度图与行为可视化)达到与专家设计课程相当的性能。本工作揭示了大模型作为具身智能系统交互式课程设计者的能力,可扩展至更广泛的进化机器人应用。
原文摘要 · Abstract (English)
Multi-task policy search is a challenging problem because policies are required to generalize beyond training cases. Curriculum learning has proven to be effective in this setting, as it introduces complexity progressively. However, designing effective curricula is labor-intensive and requires extensive domain expertise. LLM-based curriculum generation has only recently emerged as a potential solution, but was limited to operate in static, offline modes without leveraging real-time feedback from the optimizer. Here we propose an interactive LLM-assisted framework for online curriculum generation, where the LLM adaptively designs training cases based on real-time feedback from the evolutionary optimization process. We investigate how different feedback modalities, ranging from numeric metrics alone to combinations with plots and behavior visualizations, influence the LLM ability to generate meaningful curricula. Through a 2D robot navigation case study, tackled with genetic programming as optimizer, we evaluate our approach against static LLM-generated curricula and expert-designed baselines. We show that interactive curriculum generation outperforms static approaches, with multimodal feedback incorporating both progression plots and behavior visualizations yielding performance competitive with expert-designed curricula. This work contributes to understanding how LLMs can serve as interactive curriculum designers for embodied AI systems, with potential extensions to broader evolutionary robotics applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。