arXiv:2509.02278cs.GRcs.AI2025-09被引 2

用歌词和声音生成有情感的3D头像动画,更自然生动。

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

  • 通过歌词与声学联合推理生成带时间戳的运动字幕作为动画先验
  • 在真实感、表现力和情感一致性上超越现有方法,提升显著
  • 适合虚拟偶像、娱乐互动等需要细腻表情变化的场景

歌唱驱动的3D头像动画是一项具有前景但挑战重重的任务,广泛应用于虚拟形象、娱乐与教育领域。与语音不同,歌唱包含更丰富的感情色彩、动态语调和基于歌词的语义,要求生成精细且时序一致的面部动作。现有语音驱动方法常导致动画过于简单、情感平淡且语义不一致。为此,我们提出Think2Sing,一种基于扩散模型的框架,利用预训练大语言模型结合歌词与声学信息,生成语义连贯、时序一致的3D头像动画。核心创新在于引入运动字幕——一种通过新型歌唱思维链推理与声学引导检索生成的辅助语义表示,包含精确时间戳和区域化动作描述,作为可解释的运动先验。我们将任务建模为运动强度预测问题,实现对脸部区域的精细控制,增强表现力建模。为此,我们构建了一个多模态歌唱数据集,包含同步视频、声学特征与运动字幕,支持多样且富有表现力的动作学习。大量实验表明,Think2Sing在真实感、表现力与情感保真度上均优于现有最优方法,同时提供灵活可控的动画编辑能力。

原文摘要 · Abstract (English)

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric-based semantics, requiring the synthesis of fine-grained, temporally coherent facial motion. Existing speech-driven approaches often produce oversimplified, emotionally flat, and semantically inconsistent results, which are insufficient for singing animation. To address this, we propose Think2Sing, a diffusion-based framework that leverages pretrained large language models to generate semantically coherent and temporally consistent 3D head animations, conditioned on both lyrics and acoustics. A key innovation is the introduction of motion subtitles, an auxiliary semantic representation derived through a novel Singing Chain-of-Thought reasoning process combined with acoustic-guided retrieval. These subtitles contain precise timestamps and region-specific motion descriptions, serving as interpretable motion priors. We frame the task as a motion intensity prediction problem, enabling finer control over facial regions and improving the modeling of expressive motion. To support this, we create a multimodal singing dataset with synchronized video, acoustic descriptors, and motion subtitles, enabling diverse and expressive motion learning. Extensive experiments show that Think2Sing outperforms state-of-the-art methods in realism, expressiveness, and emotional fidelity, while also offering flexible, user-controllable animation editing.

3D动画歌唱生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。