用预训练数据隐含偏好,零成本提升多智能体运动生成真实感。
Direct Post-Training Preference Alignment for Multi-Agent Motion Generation Models Using Implicit Feedback from Pre-training Demonstrations
- 基于预训练示范中的隐含偏好构建生成样本间排序
- 100万参数模型生成效果媲美顶尖模仿学习大模型
- 无需人工标注偏好,降低计算与人力成本
大语言模型在具身应用中的运动生成中取得进展。尽管基于自回归的运动生成模型具备可扩展性,但其基于标记预测的目标与人类偏好存在偏差,导致生成行为偏离人类期望,因此后训练阶段的偏好对齐至关重要。然而,现有方法依赖大量人工标注的运动偏好排序,成本高昂,尤其在多智能体场景下更难实现。近期研究尝试利用预训练示范数据自动构建偏好数据,但多采用对抗假设,将所有模型生成样本视为不受欢迎,忽略了样本间相对偏好信号,降低了对齐效果并可能引发误对齐。本文提出新方法,不将所有生成样本视为同等差,而是从预训练示范中提取隐含偏好,构建模型自身生成结果间的偏好排序,提供更精细的对齐指导,且无需额外人工标注。我们在大规模交通仿真中验证该方法,结果显示,仅依赖预训练示范的隐含反馈,一个轻量级100万参数模型即可达到当前最优模仿学习大模型的表现,显著提升生成行为的真实感,同时避免高计算开销和人工标注成本。
原文摘要 · Abstract (English)
Recent advancements in LLMs have revolutionized motion generation models in embodied applications. While LLM-type auto-regressive motion generation models benefit from training scalability, there remains a discrepancy between their token prediction objectives and human preferences. As a result, models pre-trained solely with token-prediction objectives often generate behaviors that deviate from what humans would prefer, making post-training preference alignment crucial for producing human-preferred motions. Unfortunately, post-training alignment requires extensive preference rankings of motions generated by the pre-trained model, which are costly to annotate, especially in multi-agent settings. Recently, there has been growing interest in leveraging pre-training demonstrations to scalably generate preference data for post-training alignment. However, these methods often adopt an adversarial assumption, treating all pre-trained model-generated samples as unpreferred examples. This adversarial approach overlooks the valuable signal provided by preference rankings among the model's own generations, ultimately reducing alignment effectiveness and potentially leading to misaligned behaviors. In this work, instead of treating all generated samples as equally bad, we leverage implicit preferences encoded in pre-training demonstrations to construct preference rankings among the pre-trained model's generations, offering more nuanced preference alignment guidance with zero human cost. We apply our approach to large-scale traffic simulation and demonstrate its effectiveness in improving the realism of pre-trained model's generated behaviors, making a lightweight 1M motion generation model comparable to SOTA large imitation-based models by relying solely on implicit feedback from pre-training demonstrations, without additional post-training human preference annotations or high computational costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。