提出MEAT模型,实现百万像素级人体多视角生成。
MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
- 用网格注意力机制降低高分辨率多视角注意力计算复杂度
- 在1024x1024下生成一致且稠密的人体多视角图像
- 适合需要高精度人体3D生成的研究与应用
多视角扩散模型在通用物体的图像到3D生成中表现优异,但在人体数据上仍难取得理想效果,主要受限于多视角注意力难以扩展至高分辨率。本文探索了百万像素级人体多视角扩散模型,提出网格注意力(mesh attention)机制,使训练可在1024x1024分辨率下进行。该方法以着装人体网格作为中心粗略几何表示,通过光栅化与投影建立跨视角坐标对应关系,显著降低多视角注意力复杂度,同时保持视角间一致性。基于此,设计网格注意力模块并结合关键点条件,构建专用于人体的多视角扩散模型MEAT。此外,论文还提供关于使用多视角人体动作视频进行扩散训练的深入见解,缓解长期存在的数据稀缺问题。大量实验表明,MEAT能有效生成高密度、一致性的百万像素级多视角人体图像,优于现有方法。
原文摘要 · Abstract (English)
Multiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling multiview attention to higher resolutions. In this paper, we explore human multiview diffusion models at the megapixel level and introduce a solution called mesh attention to enable training at 1024x1024 resolution. Using a clothed human mesh as a central coarse geometric representation, the proposed mesh attention leverages rasterization and projection to establish direct cross-view coordinate correspondences. This approach significantly reduces the complexity of multiview attention while maintaining cross-view consistency. Building on this foundation, we devise a mesh attention block and combine it with keypoint conditioning to create our human-specific multiview diffusion model, MEAT. In addition, we present valuable insights into applying multiview human motion videos for diffusion training, addressing the longstanding issue of data scarcity. Extensive experiments show that MEAT effectively generates dense, consistent multiview human images at the megapixel level, outperforming existing multiview diffusion methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。