arXiv:2603.18588cs.CVcs.MM2026-03

用解剖合理的人体描述生成更精细的面部表情,避免肌肉冲突问题。

A Novel FACS-Aligned Anatomical Text Description Paradigm for Fine-Grained Facial Behavior Synthesis

  • 将FACS肌肉规则转化为自然语言控制信号,确保动作生理合理。
  • 构建30万+样本的图文数据集,支持细粒度面部行为合成。
  • 提出新评估指标与模型,显著提升复杂冲突场景下的真实感。

面部行为是人类非语言交流的主要媒介。现有合成方法主要采用粗粒度情绪标签或基于面部动作编码系统(FACS)的一维动作单元(AU)向量,均难以可靠生成细粒度面部行为,且易产生由冲突AU引起的解剖不合理伪影。为此,我们提出一种新任务范式:基于FACS的解剖学基础文本描述进行面部行为合成。该范式显式编码了FACS定义的肌肉运动规则、跨AU交互关系及冲突消解机制作为自然语言控制信号。为推动系统研究,我们开发了动态AU文本处理器,一个基于FACS规则的模块,可将原始AU标注转换为解剖一致的自然语言描述。利用该处理器,构建了首个大规模文本-图像配对数据集BP4D-AUText,包含超过30.2万张高质量样本。由于现有通用语义一致性指标无法捕捉解剖描述与合成肌肉运动之间的对齐,我们提出了任务专用指标AAAD(AU概率分布对齐准确率),用于量化语义一致性。最后,设计了集成解剖先验与渐进式跨模态对齐的VQ-AUFace基准框架以验证该范式。大量定量实验与用户研究证明,该范式显著优于当前最先进方法,尤其在具有挑战性的冲突AU场景下,实现了更高的解剖保真度、语义一致性和视觉质量。

原文摘要 · Abstract (English)

Facial behavior constitutes the primary medium of human nonverbal communication. Existing synthesis methods predominantly follow two paradigms: coarse emotion category labels or one-hot Action Unit (AU) vectors from the Facial Action Coding System (FACS). Neither paradigm reliably renders fine-grained facial behaviors nor resolves anatomically implausible artifacts caused by conflicting AUs. Therefore, we propose a novel task paradigm: anatomically grounded facial behavior synthesis from FACS-based AU descriptions. This paradigm explicitly encodes FACS-defined muscle movement rules, inter-AU interactions, and conflict resolution mechanisms into natural language control signals. To enable systematic research, we develop a dynamic AU text processor, a FACS rule-based module that converts raw AU annotations into anatomically consistent natural language descriptions. Using this processor, we construct BP4D-AUText, the first large-scale text-image paired dataset for fine-grained facial behavior synthesis, comprising over 302K high-quality samples. Given that existing general semantic consistency metrics cannot capture the alignment between anatomical facial descriptions and synthesized muscle movements, we propose the Alignment Accuracy of AU Probability Distributions (AAAD), a task-specific metric that quantifies semantic consistency. Finally, we design VQ-AUFace, a robust baseline framework incorporating anatomical priors and progressive cross-modal alignment, to validate the paradigm. Extensive quantitative experiments and user studies demonstrate the paradigm significantly outperforms state-of-the-art methods, particularly in challenging conflicting AU scenarios, achieving superior anatomical fidelity, semantic consistency, and visual quality.

面部生成解剖建模FACS文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。