arXiv:2604.22865cs.CV2026-04中稿 · CVPR被引 1

单张图生成可动画高精度3D人脸,无需优化

MeshLAM: Feed-Forward One-Shot Animatable Textured Mesh Avatar Reconstruction

论文配图:MeshLAM: Feed-Forward One-Shot Animatable Textured Mesh Avatar Reconstruction
图 1 · 摘自论文原文
  • 共享Transformer骨干提取特征,双路处理形状与纹理
  • 单次前向传播完成网格重建,误差仅1.78mm
  • 适合实时虚拟角色生成,对算力要求低

我们提出MeshLAM,一种前向传播的单张图像可动画化网格头像重建框架,能从单张图像生成高保真、可动画的3D头像。不同于依赖耗时测试优化或多视角数据的方法,本方法在一次前向传播中生成完整网格表示,并具备内在可动画性。采用双形状-纹理图架构,通过共享Transformer骨干提取的图像特征同步处理网格顶点与纹理图,实现一致的形体雕刻与外观建模。为防止前向变形中的网格坍缩并确保拓扑完整性,提出基于迭代GRU的解码机制,实现渐进式几何变形与纹理精炼,并引入基于重投影的纹理引导机制,将外观学习锚定于输入图像。大量实验表明,本方法在重建质量、动画能力与计算效率上均优于现有最佳方法。

原文摘要 · Abstract (English)

We introduce MeshLAM, a feed-forward framework for one-shot animatable mesh head reconstruction that generates high-fidelity, animatable 3D head avatars from a single image. Unlike previous work that relies on time-consuming test-time optimization or extensive multi-view data, our method produces complete mesh representations with inherent animatability from a single image in a single forward pass. Our approach employs a dual shape and texture map architecture that simultaneously processes mesh vertices and texture map with extracted image features from a shared transformer backbone, allowing for coherent shape carving and appearance modeling. To prevent mesh collapse and ensure topological integrity during feed-forward deformation, we propose an iterative GRU-based decoding mechanism with progressive geometry deformation and texture refinement, coupled with a novel reprojection-based texture guidance mechanism that anchors appearance learning to the input image. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches in reconstruction quality, animation capability, and computational efficiency. Project page at https://meshlam.github.io.

3D人脸单图重建可动画网格生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。