arXiv:2609.04903cs.CV2026-09

让两人合唱时的互动更自然,通过动态信号控制表情与动作协调。

InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond

论文配图:InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond
图 1 · 摘自论文原文
  • 用交互逻辑值建模两人在音乐关键点的互动强度。
  • 生成动画比现有方法更协调且符合音乐节奏,保持个人风格。
  • 适用于多人合唱,可灵活控制互动时机与程度。

我们提出 InterSing,一个用于生成双人合唱表演中真实3D头部动画的框架。与独唱不同,双人表演需在音乐显著时刻(如乐句边界、同步节奏、问答段落)实现个体表现力与间歇性互动的平衡。由于这些互动稀疏且依赖节奏,现有音频驱动动画与对话模型难以捕捉其结构。我们的核心洞察是:双人协作可表示为随时间变化的信号,反映表演者之间互动强度。基于此,我们引入交互逻辑值(interaction logits),一种可解释的潜在表征,用于建模每个时间步的跨表演者参与度。通过弱监督学习该逻辑值,并将其作为条件输入到联合音频特征与互动动态的交互感知扩散模型中。该方法支持统一的多模式生成,涵盖独立运动、协调行为及平滑过渡。实验表明,InterSing 生成的动画更具表现力且协调性更强,音乐对齐更精准,同时保留每位表演者的特征动作风格。我们还证明该框架可推广至多表演者场景,并提供对互动时机与方式的直观控制。

原文摘要 · Abstract (English)

We present InterSing, a framework for generating realistic 3D head animations for duet singing performances. Unlike solo singing, duet performance requires each singer to balance individual expressiveness with intermittent interaction at musically salient moments, such as phrase boundaries, synchronized rhythms, and call-and-response passages. Because these interactions are sparse and rhythm-dependent, existing audio-driven animation methods and conversational interaction models do not adequately capture their structure. Our key insight is that duet coordination can be represented as a time-varying signal that reflects how strongly performers engage with one another throughout a song. Based on this observation, we introduce interaction logits, an interpretable latent representation that models the degree of cross-performer engagement at each time step. We learn these logits using weak supervision and use them to condition an interaction-aware diffusion model jointly driven by audio features and interaction dynamics. This formulation enables unified multi-mode generation, spanning independent motion, coordinated behavior, and smooth transitions between them. Experiments show that InterSing generates realistic and expressive singing head animations with stronger coordination and musical alignment than existing methods, while preserving each performer's characteristic motion style. We further demonstrate that the same formulation generalizes to multi-singer performances and provides intuitive control over when and how performers engage.

3D动画语音驱动双人互动扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。