arXiv:2603.25020cs.CV2026-03

让虚拟听众更生动,解决表情僵化问题。

GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization

  • 用自回归流匹配实现稳定训练,提升生成质量。
  • 通过分组奖励解耦优化,显著提高动作多样性。
  • 支持文本语义控制,可定制化表达反应。

在虚拟人类合成中,生成两人互动时的逼真3D头部动作是一项重大挑战。现有方法虽在说话人动作上表现优异,但听者动作常出现‘回归均值’现象,导致面部僵化,并缺乏复杂非语言动作的参数空间。本文提出GDPO-Listener框架,实现高表现力的说话与倾听动作生成。首先,引入自回归流匹配架构,支持稳定的监督学习;其次,为克服运动僵硬,采用分组奖励解耦策略(GDPO),通过在不同FLAME参数组间解耦奖励归一化,显式激励高方差的表达性生成;最后,实现对语义文本的显式控制,支持个性化响应。在Seamless Interaction和DualTalk数据集上的大量评估表明,该方法在长期运动方差、视觉表现力和语义可控性方面均优于现有基线。

原文摘要 · Abstract (English)

Generating realistic 3D head motion for dyadic interactions is a significant challenge in virtual human synthesis. While recent methods achieve impressive results with speaking heads, they frequently suffer from the `Regression-to-the-Mean' problem in listener motions, collapsing into static faces, and lack the parameter space for complex nonverbal motions. In this paper, we propose GDPO-Listener, a novel framework that achieves highly expressive speaking and listening motion generation. First, we introduce an Auto-Regressive Flow Matching architecture enabling stable supervised learning. Second, to overcome kinematic stillness, we apply the Group reward-Decoupled Policy Optimization (GDPO). By isolating reward normalization across distinct FLAME parameter groups, GDPO explicitly incentivizes high variance expressive generations. Finally, we enable explicit semantic text control for customizable responses. Extensive evaluations across the Seamless Interaction and DualTalk datasets demonstrate superior performance compared to existing baselines on long-term kinematic variance, visual expressivity and semantic controllability.

3D生成交互虚拟人动作生成文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。