构建10万组人脸表情与机械控制数据,实现真人表情精准复刻
X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation
- 分两阶段建模视觉特征与机械控制的映射关系
- 在10万组数据上实现高保真表情迁移,物理验证效果优
- 适合机器人表情生成、人机交互研究者参考
从人类到类人机器人精细面部表情迁移面临独特模式识别挑战,源于生物面部动态与机械控制空间间显著域差距。尽管对话头视觉合成已快速进展,但将高维视觉线索映射为精确、物理受限的驱动信号仍是开放问题,主要因缺乏大规模配对数据。为此,我们提出X2C,一个包含10万组(图像,控制值)配对的综合性基准数据集。不同于现有资源,X2C包含30个连续控制参数标注的细微、物理真实表情,确立了该任务的高保真标准。基于此资源,我们提出X2CNet,一种两阶段深度学习框架,显式解耦视觉运动特征与机械控制回归,建模人类感知线索与类人机器人驱动间的对应关系。大量实验,包括定量基准测试与真实世界物理验证,表明该方法在跨域一致性上表现更优,支持鲁棒的野外表情模仿。代码与数据:https://lipzh5.github.io/X2CNet/
原文摘要 · Abstract (English)
Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significant domain gap between biological facial dynamics and mechanical control spaces. While visual synthesis of talking heads has advanced rapidly, mapping high-dimensional visual cues to precise, physically constrained actuation signals remains an open problem, primarily due to the lack of large-scale paired data. To bridge this gap, we introduce X2C, a comprehensive benchmark dataset comprising 100,000 (image, control value) pairs. Unlike existing resources, X2C features nuanced, physically grounded expressions annotated with 30 continuous control parameters, establishing a high-fidelity standard for this task. Building on this resource, we propose X2CNet, a two-stage deep learning framework that explicitly decouples visual motion features from mechanical control regression to model the correspondence between human perceptual cues and humanoid actuation. Extensive experiments, including quantitative benchmarking and real-world physical validation, demonstrate that our approach achieves superior cross-domain consistency and enables robust, in-the-wild expression imitation. Code and Data: https://lipzh5.github.io/X2CNet/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。