用人类反馈生成更自然的双人互动表情,避免身份偏见。
Facial Expression Generation Aligned with Human Preference for Natural Dyadic Interaction
- 将表情生成视为无身份偏差的动作学习,利用人类反馈评估
- 在两个基准上表现优于现有方法,表情更符合人类偏好
- 适合研究人机交互、情感计算及虚拟角色生成的学者
实现自然的双人互动需要生成情感恰当且社会契合的表情。人类反馈为引导这种对齐提供了有效机制,但如何将其融入表情生成仍待探索。本文提出一种基于人类反馈的面部表情生成方法,通过框架化身份无关的表情生成为动作学习过程,使人类反馈能无视觉或身份偏见地评估表达有效性。我们建立了一个闭环反馈系统,使听者表情动态响应说话者对话线索的变化。具体而言,通过监督微调训练一个视觉-语言-动作模型,将说话者的多模态信号映射为3D可变形模型的可控低维表达表示。进一步引入人类反馈强化学习策略,结合高质量表达响应模仿与批评者引导优化。在两个基准上的实验表明,该方法能有效对齐人类偏好,并取得更优性能。
原文摘要 · Abstract (English)
Achieving natural dyadic interaction requires generating facial expressions that are emotionally appropriate and socially aligned with human preference. Human feedback offers a compelling mechanism to guide such alignment, yet how to effectively incorporate this feedback into facial expression generation remains underexplored. In this paper, we propose a facial expression generation method aligned with human preference by leveraging human feedback to produce contextually and emotionally appropriate expressions for natural dyadic interaction. A key to our method is framing the generation of identity-independent facial expressions as an action learning process, allowing human feedback to assess their validity free from visual or identity bias. We establish a closed feedback loop in which listener expressions dynamically respond to evolving conversational cues of the speaker. Concretely, we train a vision-language-action model via supervised fine-tuning to map the speaker's multimodal signals into controllable low-dimensional expression representations of a 3D morphable model. We further introduce a human-feedback reinforcement learning strategy that integrates the imitation of high-quality expression response with critic-guided optimization. Experiments on two benchmarks demonstrate that our method effectively aligns facial expressions with human preference and achieves superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。