用扩散模型生成自然同步的听者反应,突破传统方法局限。
DiffListener: Discrete Diffusion Model for Listener Generation
- 基于离散扩散模型,非自回归生成听者头部反应。
- 融合面部动态差分信息,实现时序一致性建模。
- 用户测试验证反应自然且与讲话者高度同步。
听者头部生成(LHG)任务旨在根据说话者的多模态线索生成自然的非语言反应。现有方法或仅依赖有限模态(如音频和面部信息),或采用自回归生成方式,存在预测误差累积等问题。为此,我们提出 DiffListener,一种基于离散扩散模型的非自回归听者头部生成方法。该模型输入说话者的面部、音频与文本信息,并引入面部差分信息以显式建模表情与动作的时序动态。通过这种显式动态建模,DiffListener 能够非自回归地生成连贯的反应序列。大量实验表明,该方法在定量与定性评估中均达到当前最优性能。用户研究显示,生成的反应自然且与讲话者高度同步。代码与演示视频见 https://siyeoljung.github.io/DiffListener。
原文摘要 · Abstract (English)
The listener head generation (LHG) task aims to generate natural nonverbal listener responses based on the speaker's multimodal cues. While prior work either rely on limited modalities (e.g. audio and facial information) or employ autoregressive approaches which have limitations such as accumulating prediction errors. To address these limitations, we propose DiffListener, a discrete diffusion based approach for non-autoregressive listener head generation. Our model takes the speaker's facial information, audio, and text as inputs, additionally incorporating facial differential information to represent the temporal dynamics of expressions and movements. With this explicit modeling of facial dynamics, DiffListener can generate coherent reaction sequences in a non-autoregressive manner. Through comprehensive experiments, DiffListener demonstrates state-of-the-art performance in both quantitative and qualitative evaluations. The user study shows that DiffListener generates natural context-aware listener reactions that are well synchronized with the speaker. The code and demo videos are available in https://siyeoljung.github.io/DiffListener
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。