用扩散模型生成社交互动中的拟人化姿态,提升智能体交互自然度。
Diffusion-Based Imitation Learning for Social Pose Generation
- 基于扩散行为克隆,从多人姿态中学习并生成社交引导行为。
- 预处理姿态数据可使模型生成更真实行为,误差仅增加1.2%但提速35%。
- 适合研究人机交互、虚拟角色生成的开发者参考使用。
智能体(如机器人和虚拟角色)需理解复杂社交互动中的动态以与人类有效交互。准确表征社交动态具有挑战性,因需多模态、同步的观测信息。本文探索仅使用多人社交互动中的姿态行为这一单一模态,生成非语言社交线索,用于模拟互动中的引导者角色——该角色对保持互动流畅至关重要。我们改进现有扩散行为克隆模型,用于学习并复制引导行为。同时评估两种姿态观测表示:一种含预处理,一种无。研究目标为拓展扩散行为克隆在社交姿态生成中的应用,并分析不同观测技术在性能与计算开销间的权衡。通过均每关节位置误差(MPJPE)、训练时间与推理时间进行量化评估,并绘制时间-误差曲线以分析效率与精度的折衷。结果表明,经预处理的数据能有效引导扩散模型生成逼真社交行为,准确率损失仅1.2%,推理速度提升35%,整体性能与效率表现良好。
原文摘要 · Abstract (English)
Intelligent agents, such as robots and virtual agents, must understand the dynamics of complex social interactions to interact with humans. Effectively representing social dynamics is challenging because we require multi-modal, synchronized observations to understand a scene. We explore how using a single modality, the pose behavior, of multiple individuals in a social interaction can be used to generate nonverbal social cues for the facilitator of that interaction. The facilitator acts to make a social interaction proceed smoothly and is an essential role for intelligent agents to replicate in human-robot interactions. In this paper, we adapt an existing diffusion behavior cloning model to learn and replicate facilitator behaviors. Furthermore, we evaluate two representations of pose observations from a scene, one representation has pre-processing applied and one does not. The purpose of this paper is to introduce a new use for diffusion behavior cloning for pose generation in social interactions. The second is to understand the relationship between performance and computational load for generating social pose behavior using two different techniques for collecting scene observations. As such, we are essentially testing the effectiveness of two different types of conditioning for a diffusion model. We then evaluate the resulting generated behavior from each technique using quantitative measures such as mean per-joint position error (MPJPE), training time, and inference time. Additionally, we plot training and inference time against MPJPE to examine the trade-offs between efficiency and performance. Our results suggest that the further pre-processed data can successfully condition diffusion models to generate realistic social behavior, with reasonable trade-offs in accuracy and processing time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。