用长文本精准生成和编辑3D人脸,细节更真实。
RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing

- 通过检索语义相关的面部几何先验,提升生成精度。
- 在局部几何准确率上优于现有方法,支持精细区域编辑。
- 适合需要高保真人脸生成与编辑的影视、游戏开发者。
由于难以将长篇描述转化为细粒度面部几何,文本驱动的3D人脸生成与编辑仍具挑战。现有方法多对齐全局语义与面部结构,但常忽略细微局部变形,如眉部紧张、脸颊收缩和不对称口部动作,导致几何保真度和编辑精度受限。为此,我们构建了FaME-G2E,一个大规模多模态数据集,包含详细的文本-网格标注及配对的文本-混合形状样本,用于统一的3D人脸生成与编辑。基于此,提出RAGMesh,一种检索增强框架,利用文本相关几何先验提升高保真人脸合成与编辑能力。其多尺度检索融合(MSRF)模块在混合形状空间中检索并融合语义一致的全局与局部先验,抑制冲突局部变形,保留连贯变形模式。此外,引入自适应RAG引导监督(AdaRAGS),一种区域感知约束,显式对齐文本语义与对应面部区域,增强区域可控性与编辑精度。在FaME-G2E上的大量实验表明,RAGMesh在局部几何精度、文本引导可控性、区域编辑精度和推理效率方面均优于当前最优方法。视频演示见https://youtu.be/Yr0_XkpWcNk,源代码与数据集将在论文接收后发布。
原文摘要 · Abstract (English)
Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geometry. Existing methods primarily align global textual semantics with facial structures but often struggle to capture subtle local deformations, such as eyebrow tension, cheek contraction, and asymmetric mouth motions, resulting in limited geometric fidelity and editing precision. To facilitate fine-grained text-driven facial modeling, we first construct FaME-G2E, a large-scale multimodal dataset containing detailed text--mesh annotations and paired text--blendshape samples for unified 3D facial generation and editing. Based on this dataset, we propose RAGMesh, a retrieval-augmented framework that leverages text-correlated geometric priors to improve high-fidelity facial synthesis and editing. Specifically, the Multi-Scale Retrieval Fusion (MSRF) module retrieves semantically consistent global and regional facial priors and fuses them in the blendshape space, suppressing conflicting local deformations while preserving coherent deformation patterns. Furthermore, we introduce Adaptive RAG-guided Supervision (AdaRAGS), a region-aware constraint that explicitly aligns textual semantics with corresponding facial regions, enhancing regional controllability and editing accuracy. Extensive experiments on FaME-G2E demonstrate that RAGMesh achieves superior performance over state-of-the-art methods in local geometric accuracy, text-guided controllability, regional editing precision, and inference efficiency. Video demo is available at https://youtu.be/Yr0_XkpWcNk, and the source code and dataset will be released upon paper acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。