用统一模型秒级重建高质量3D avatar,支持多源输入
FastAvatar: Towards Unified and Fast 3D Avatar Reconstruction with Large Gaussian Reconstruction Transformers
- 基于大视觉-几何融合的Transformer架构,统一处理单图到视频输入
- 在1秒内完成重建,比现有方法快数倍且质量更高
- 适合需要快速生成3D数字人场景的应用,如虚拟会议、游戏
尽管3D数字人重建已取得显著进展,但仍面临计算复杂度高、对数据质量敏感、数据利用率低等挑战。本文提出FastAvatar,一种前馈式3D数字人框架,可灵活利用日常记录(如单张图像、多视角观测或单目视频)在数秒内重构出高质量3D高斯溅射(3DGS)模型,仅需一个统一模型。其核心是大型高斯重建变压器(LGRT),包含三项关键设计:第一,3DGS Transformer聚合多帧信息并注入初始3D提示,预测对应注册的规范3DGS表示;第二,多粒度引导编码(相机位姿、表情系数、头部姿态),缓解不同长度输入带来的动画错位问题;第三,通过特征点追踪与分片融合损失实现增量高斯聚合。结合上述设计,FastAvatar支持增量重建——随着输入增多逐步提升质量,不浪费数据。这实现了质量-速度可调的高效3D数字人建模范式。大量实验表明,FastAvatar在质量与速度上均优于现有方法。
原文摘要 · Abstract (English)
Despite significant progress in 3D avatar reconstruction, it still faces challenges such as high time complexity, sensitivity to data quality, and low data utilization. We propose FastAvatar, a feedforward 3D avatar framework capable of flexibly leveraging diverse daily recordings (e.g., a single image, multi-view observations, or monocular video) to reconstruct a high-quality 3D Gaussian Splatting (3DGS) model within seconds, using only a single unified model. The core of FastAvatar is a Large Gaussian Reconstruction Transformer (LGRT) featuring three key designs: First, a 3DGS transformer aggregating multi-frame cues while injecting initial 3D prompt to predict the corresponding registered canonical 3DGS representations; Second, multi-granular guidance encoding (camera pose, expression coefficient, head pose) mitigating animation-induced misalignment for variable-length inputs; Third, incremental Gaussian aggregation via landmark tracking and sliced fusion losses. Integrating these features, FastAvatar enables incremental reconstruction, i.e., improving quality with more observations without wasting input data as in previous works. This yields a quality-speed-tunable paradigm for highly usable 3D avatar modeling. Extensive experiments show that FastAvatar has a higher quality and highly competitive speed compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。