用大规模数据训练高保真3D虚拟人,兼顾泛化与细节控制。
Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining
- 先在百万级野外视频上预训练,再用高质量数据微调
- 支持发型、服装、种族等广泛泛化,面部表情精细到手指
- 无需额外训练即能处理光照变化和松散衣物,适合虚拟人应用
高质量3D虚拟人建模面临保真度与泛化能力的权衡。多视角工作室数据可实现高保真建模并精确控制表情与姿态,但因规模有限且与真实场景存在域差距,难以泛化;而基于百万级野外样本训练的大规模虚拟人模型虽具备强泛化能力,却因3D固有歧义导致质量偏低。为此,我们提出大型编码器虚拟人(LCA),一种高保真、全身3D虚拟人模型,可在前馈模式下高效泛化至全球规模人群。受大语言模型与视觉基础模型启发,首次在3D虚拟人领域引入大规模预/微调范式:在100万条野外视频上预训练以学习广泛的外观与几何先验,再在高质量标注数据上微调以提升表现力与保真度。LCA在发型、服装、人种等多样性上表现出色,支持精确的细粒度面部表情与指节级动作控制,并保持强身份一致性。值得注意的是,尽管未直接监督,其仍展现出对光照重演和松散衣物的涌现泛化能力,以及对风格化图像的零样本鲁棒性。
原文摘要 · Abstract (English)
High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and the domain gap between the studio environment and the real world. On the other hand, recent large-scale avatar models trained on millions of in-the-wild samples show promise for generalization across a wide range of identities, yet the resulting avatars are often of low-quality due to inherent 3D ambiguities. To address this, we present Large-Scale Codec Avatars (LCA), a high-fidelity, full-body 3D avatar model that generalizes to world-scale populations in a feedforward manner, enabling efficient inference. Inspired by the success of large language models and vision foundation models, we present, for the first time, a pre/post-training paradigm for 3D avatar modeling at scale: we pretrain on 1M in-the-wild videos to learn broad priors over appearance and geometry, then post-train on high-quality curated data to enhance expressivity and fidelity. LCA generalizes across hair styles, clothing, and demographics while providing precise, fine-grained facial expressions and finger-level articulation control, with strong identity preservation. Notably, we observe emergent generalization to relightability and loose garment support to unconstrained inputs, and zero-shot robustness to stylized imagery, despite the absence of direct supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。