arXiv:2607.10313cs.GRcs.SD2026-07

用单人说话动作先验,让两人对话更自然流畅。

Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic Priors

论文配图:Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic Priors
图 1 · 摘自论文原文
  • 以单人动作模型为基础,分离说话动作与互动影响。
  • 在DualTalk上实现最佳量化与主观评价表现。
  • 适合需要高效生成对话数字人的研究者使用。

双人对话动作生成对打造真实互动数字人至关重要。现有方法多采用统一的双人生成器,但往往将自我发声驱动的动作与对方回应的社会反馈耦合,导致互动特异性成分隐含且未充分利用预训练单人动作模型已学习的语音-动作对应关系。本文提出Learn2Chat框架,将双人动作建模为预训练单人动作先验上的互动调制。该设计分离了内在语音驱动动作与社会互动效应,实现更结构化的互动建模。具体地,引入基于单人锚定的动作分解机制,利用单人数据学习到的语义动作流形,解耦语音驱动的动作动态与互动诱导的调制,从双人序列中提取清晰的互动表征。在此表征空间上,通过跨分支注意力与互动对齐,设计交叉注意力互动隐变量预测模块,将配对语音信号映射为互动隐变量。推理时,预测的互动隐变量调制标准单人动作,以数据高效方式生成连贯同步的双人行为。在DualTalk基准上的大量实验表明,Learn2Chat在量化指标与感知评估上均达到当前最优。此外,该框架具备模型无关性,可无缝集成多种预训练单人动作主干,凸显先验复用与互动适配在可扩展对话动作生成中的有效性。

原文摘要 · Abstract (English)

Dyadic conversational motion generation is essential for realistic interactive digital humans. Existing approaches typically model conversational behaviors within unified dyadic generators. However, such holistic formulations tend to couple self-speech-driven motion with partner-responsive social feedback, leaving the interaction-specific component implicit and underutilizing the speech-motion correspondence already learned by pretrained monologic motion models. We propose Learn2Chat, a unified framework that models dyadic motion as interaction modulation over pretrained monologic motion priors. This design separates intrinsic speech-driven motion from social interaction effects and enables more structured interaction modeling. Specifically, we introduce a Monologic-Anchored Motion Factorization scheme that leverages the semantic motion manifold learned from monologic data to disentangle audio-driven motion dynamics from interaction-induced modulation, yielding clean interaction representations from dyadic sequences. On top of this representation space, a Cross-Attentive Interaction Latent Prediction module maps paired speech signals to interaction latents through cross-branch attention and interaction alignment. During inference, the predicted interaction latents modulate canonical monologic motion to generate coherent and synchronized dyadic behaviors in a data-efficient manner. Extensive experiments on the DualTalk benchmark demonstrate that Learn2Chat achieves state-of-the-art performance across both quantitative metrics and perceptual evaluations. Moreover, the framework is model-agnostic and seamlessly integrates with diverse pretrained monologic motion backbones, highlighting the effectiveness of prior reuse and interaction adaptation for scalable conversational motion generation. More visual results are available on the project page.

对话生成动作建模先验复用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。