arXiv:2502.20387cs.CV2025-02CVPR被引 20

仅用几秒视频即可快速生成逼真3D虚拟说话人。

InsTaG: Learning Personalized 3D Talking Head from Few-Second Video

  • 用轻量级3DGS模型+通用运动先验,快速建模新身份。
  • 仅需少量数据,即可生成高质量、高个性化的3D说话头像。
  • 适合需要快速生成个性化虚拟形象的场景,如直播、虚拟助手。

尽管基于辐射场的方法在生成逼真个性化3D说话头方面表现优异,但其对训练数据量和时间要求过高。本文提出InsTaG,一种可从极少训练数据中快速学习真实个性化3D说话头的框架。该框架基于轻量级3DGS个人化合成器与通用运动先验,实现高质量且高效的个性化适配。首先,提出无身份预训练策略,使个人化模型可在长视频语料库上进行预训练,并提取通用运动先验。随后,设计运动对齐适配策略,在少量数据下自适应对齐目标头部至预训练场,并约束鲁棒的动态头部结构。实验表明,该方法在多种数据场景下均能高效生成高质量个性化3D说话头。

原文摘要 · Abstract (English)

Despite exhibiting impressive performance in synthesizing lifelike personalized 3D talking heads, prevailing methods based on radiance fields suffer from high demands for training data and time for each new identity. This paper introduces InsTaG, a 3D talking head synthesis framework that allows a fast learning of realistic personalized 3D talking head from few training data. Built upon a lightweight 3DGS person-specific synthesizer with universal motion priors, InsTaG achieves high-quality and fast adaptation while preserving high-level personalization and efficiency. As preparation, we first propose an Identity-Free Pre-training strategy that enables the pre-training of the person-specific model and encourages the collection of universal motion priors from long-video data corpus. To fully exploit the universal motion priors to learn an unseen new identity, we then present a Motion-Aligned Adaptation strategy to adaptively align the target head to the pre-trained field, and constrain a robust dynamic head structure under few training data. Experiments demonstrate our outstanding performance and efficiency under various data scenarios to render high-quality personalized talking heads.

3D生成说话头轻量化快速建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。