用视频参考生成更像真人的虚拟人视频,能复现说话节奏和微表情。
Avatar V: Scaling Video-Reference Avatar Video Generation

- 直接用参考视频的完整帧序列建模身份与动态行为,不依赖静态图片
- 生成1080p无限时长视频,多项指标领先现有系统,人类评测也胜出
- 适合需要高保真人像生成的影视、虚拟主播等场景
生成不仅外观相似且行为可识别的虚拟人视频仍面临挑战:现有方法多依赖单张静态图像,信息不足,无法捕捉动态特征;而常规像素级损失未能关注决定形象真实性的关键面部区域。本文提出Avatar V,一个面向生产的规模化框架,通过视频参考条件化身份建模解决上述问题。模型不再将身份压缩为固定长度嵌入,而是直接以参考视频的完整标记序列作为输入,通过注意力机制学习同时还原静态身份属性(面部结构、皮肤纹理)与动态行为模式(说话节奏、微表情)。引入稀疏参考注意力机制,实现任意长视频的线性复杂度建模;设计运动表征流支持闭环说话风格迁移;并采用身份感知超分辨率精修模块,继承完整参考条件。训练依托100M+训练片段(来自50M原始视频),采用五阶段流程:光流匹配预训练、个性微调、双阶段蒸馏(加速超10倍)、以及基于强化学习的人类偏好对齐,在数千个GPU上部署。该系统可生成1080p无限时长视频,在跨场景基准测试中,身份保留、口型同步与生成质量均达顶尖水平,显著优于Seedance 2.0、Kling O3 Pro、Veo 3.1与OmniHuman 1.5,自动化评估与人工评测均表现优异。
原文摘要 · Abstract (English)
Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, faithfully reproducing their talking rhythm, gestural tendencies, and expression dynamics, remains an open challenge. Existing methods predominantly condition on single static images, which provide insufficient identity information and cannot capture dynamic motion traits, while standard pixel-level objectives underserve the perceptually critical facial regions that determine avatar fidelity. We present Avatar V, a production-scale framework that addresses these limitations through video-reference-conditioned identity modeling. Rather than compressing identity into fixed-size embeddings, the model conditions directly on the full token sequence of a reference video, learning to reproduce both static identity attributes (facial geometry, skin texture) and dynamic behavioral patterns (talking rhythm, micro-expressions) through attention over the reference context. We introduce Sparse Reference Attention, an asymmetric mechanism achieving linear-complexity conditioning on arbitrarily long references; a motion representation stream enabling closed-loop talking style transfer; and an identity-aware super-resolution refiner inheriting the full reference conditioning. These are supported by a data engine curating 100M+ training clips from 50M raw videos, and a five-stage training pipeline with flow matching pre-training, personality fine-tuning, two-phase distillation (>10x acceleration), and RLHF alignment, deployed across thousands of GPUs. Avatar V generates 1080p videos of unlimited duration, achieving state-of-the-art identity preservation, lip synchronization, and generation quality on our cross-scene benchmark, consistently outperforming leading systems including Seedance 2.0, Kling O3 Pro, Veo 3.1, and OmniHuman 1.5 in both automated metrics and human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。