arXiv:2507.05906cs.LGcs.AI2025-07被引 4

对比两种演示学习方法,指导如何根据任务需求选择。

Feature-Based vs. GAN-Based Learning from Demonstrations: When and Why

  • 比较特征法与GAN法在奖励函数结构上的差异
  • 指出特征法适合高精度模仿,GAN法更易扩展适应
  • 强调应按任务目标如精度、多样性选方法

本文对基于特征和基于GAN的演示学习方法进行对比分析,重点考察奖励函数结构及其对策略学习的影响。特征法提供密集且可解释的奖励,擅长高保真动作模仿,但需复杂参考表示,在非结构化环境中泛化能力差。GAN法采用隐式分布监督,具备良好可扩展性和适应灵活性,但训练不稳定,奖励信号粗糙。近期进展均指向结构化运动表征的重要性,能实现平滑过渡、可控生成和更好任务整合。我们认为两类方法的差异日益细化:并非一方取代另一方,而是应根据任务优先级如保真度、多样性、可解释性与适应性做出选择。本文梳理了算法权衡与设计考量,为演示学习中的方法选择提供系统性决策框架。

原文摘要 · Abstract (English)

This survey provides a comparative analysis of feature-based and GAN-based approaches to learning from demonstrations, with a focus on the structure of reward functions and their implications for policy learning. Feature-based methods offer dense, interpretable rewards that excel at high-fidelity motion imitation, yet often require sophisticated representations of references and struggle with generalization in unstructured settings. GAN-based methods, in contrast, use implicit, distributional supervision that enables scalability and adaptation flexibility, but are prone to training instability and coarse reward signals. Recent advancements in both paradigms converge on the importance of structured motion representations, which enable smoother transitions, controllable synthesis, and improved task integration. We argue that the dichotomy between feature-based and GAN-based methods is increasingly nuanced: rather than one paradigm dominating the other, the choice should be guided by task-specific priorities such as fidelity, diversity, interpretability, and adaptability. This work outlines the algorithmic trade-offs and design considerations that underlie method selection, offering a framework for principled decision-making in learning from demonstrations.

演示学习奖励函数方法对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。