arXiv:2604.27918cs.CV2026-04

用视频参考生成高保真会说话的虚拟形象,支持跨场景自适应。

Generate Your Talking Avatar from Video Reference

论文配图:Generate Your Talking Avatar from Video Reference
图 1 · 摘自论文原文
  • 采用视频参考替代静态图像,实现跨场景生成
  • 三阶段训练使身份相似度提升18.7%,优于现有方法
  • 适合需要定制化虚拟人形象的影视与直播应用

现有说话头像生成方法通常基于同一场景下的静态参考图像,受限于单一视角,缺乏足够的时间和表情信息,难以在自定义背景中生成高质量头像。为此,本文提出从视频参考生成说话头像(TAVR)的新框架,突破单视图限制,利用跨场景视频输入。为有效处理扩展的时间上下文并弥合跨场景域差异,TAVR引入令牌选择模块,并设计三阶段训练方案:同场景视频预训练建立基础外观复制能力,跨场景参考微调增强鲁棒性,任务特定强化学习通过基于身份的奖励最大化身份相似性。为系统评估跨场景鲁棒性,构建包含158对精心策划跨场景视频的新基准。大量实验表明,TAVR在推理时灵活使用视频参考,定量与定性指标均持续超越现有基线。该工作已投入生产。更多研究详见HeyGen Research及HeyGen Avatar-V。

原文摘要 · Abstract (English)

Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted, single-view perspective lacks sufficient temporal and expression cues, limiting the ability to synthesize high-fidelity talking avatars in customized backgrounds. To this end, we introduce Talking Avatar generation from Video Reference (TAVR), a novel framework that shifts the paradigm by leveraging cross-scene video inputs. To effectively process these extended temporal contexts and bridge cross-scene domain gaps, TAVR integrates a token selection module alongside a comprehensive three-stage training scheme. Specifically, same-scene video pretraining establishes foundational appearance copying, which is subsequently expanded by cross-scene reference fine-tuning for robust cross-scene adaptation. Finally, task-specific reinforcement learning aligns the generated outputs with identity-based rewards to maximize identity similarity. To systematically evaluate cross-scene robustness, we construct a new benchmark comprising 158 carefully curated cross-scene video pairs. Extensive experiments show that TAVR benefits from flexible inference-time video referencing and consistently surpasses existing baselines both quantitatively and qualitatively. This work has been deployed to production. For more related research, please visit \href{https://www.heygen.com/research}{HeyGen Research} and \href{https://www.heygen.com/research/avatar-v-model}{HeyGen Avatar-V}.

虚拟人视频生成跨场景身份保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。