arXiv:2602.23165cs.CV2026-02被引 5

DyaDiT让数字人对话时手势更自然,能理解两人互动中的社交动态。

DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation

  • 基于双人语音与社交上下文,融合双方信息生成互动手势
  • 在标准指标上超越现有方法,用户偏好度显著提升
  • 适合虚拟助手、元宇宙角色等需要自然人际交互的场景

生成逼真的对话手势对于实现数字人之间自然且富有社交性的交互至关重要。然而,现有方法通常仅将单路音频映射为单一说话者的动作,未考虑社交背景或两人之间的相互作用。我们提出 DyaDiT,一种多模态扩散变压器,可从双人对话音频中生成符合语境的人体动作。在 Seamless Interaction Dataset 上训练,DyaDiT 接收双人音频及可选的社交上下文标记,生成符合情境的动作。它融合双方面信息以捕捉互动动态,使用动作字典编码动作先验,并可选择性利用对话伙伴的手势以生成更具响应性的动作。我们在标准动作生成指标上进行评估,并开展定量用户研究,结果表明 DyaDiT 不仅在客观指标上优于现有方法,且用户偏好度显著更高,凸显其鲁棒性和社交友好性。代码与模型将在论文录用后发布。

原文摘要 · Abstract (English)

Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without considering social context or modeling the mutual dynamics between two people engaging in conversation. We present DyaDiT, a multi-modal diffusion transformer that generates contextually appropriate human motion from dyadic audio signals. Trained on Seamless Interaction Dataset, DyaDiT takes dyadic audio with optional social-context tokens to produce context-appropriate motion. It fuses information from both speakers to capture interaction dynamics, uses a motion dictionary to encode motion priors, and can optionally utilize the conversational partner's gestures to produce more responsive motion. We evaluate DyaDiT on standard motion generation metrics and conduct quantitative user studies, demonstrating that it not only surpasses existing methods on objective metrics but is also strongly preferred by users, highlighting its robustness and socially favorable motion generation. Code and models will be released upon acceptance.

手势生成多模态扩散模型社交交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。