用关键帧与过渡帧稀疏表示面部动作,提升听觉交互的生成质量。
When Less Is More: A Sparse Facial Motion Structure For Listening Motion Learning
- 将连续面部动作压缩为关键帧+过渡帧的稀疏序列
- 生成动作更保时序结构,且个体差异性更强
- 适合需要高真实感的对话机器人行为建模
有效的人类行为建模对人机交互至关重要。当前预测双人对话中倾听头部行为的先进方法采用连续到离散的表征,将连续面部动作序列转换为离散潜在标记。然而,非语言面部动作因时间变化性和多模态性带来独特挑战,现有离散动作标记表征难以捕捉底层非语言面部模式,导致倾听头部生成训练效率低、保真度差。本研究提出一种新方法,通过将长序列编码为关键帧与过渡帧的稀疏序列,识别重要动作步骤并插值中间帧,在保持动作时序结构的同时增强学习过程中的实例多样性。此外,将该稀疏表征应用于倾听头部预测任务,证明其有助于改善面部动作模式的解释能力。
原文摘要 · Abstract (English)
Effective human behavior modeling is critical for successful human-robot interaction. Current state-of-the-art approaches for predicting listening head behavior during dyadic conversations employ continuous-to-discrete representations, where continuous facial motion sequence is converted into discrete latent tokens. However, non-verbal facial motion presents unique challenges owing to its temporal variance and multi-modal nature. State-of-the-art discrete motion token representation struggles to capture underlying non-verbal facial patterns making training the listening head inefficient with low-fidelity generated motion. This study proposes a novel method for representing and predicting non-verbal facial motion by encoding long sequences into a sparse sequence of keyframes and transition frames. By identifying crucial motion steps and interpolating intermediate frames, our method preserves the temporal structure of motion while enhancing instance-wise diversity during the learning process. Additionally, we apply this novel sparse representation to the task of listening head prediction, demonstrating its contribution to improving the explanation of facial motion patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。