自动为不同脸型生成可制造的机械面部,让机器人能自然对话
Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots

- 用参数化连杆模板和分层算法,从2D图自动生成3D机械结构
- 合成机制在多种脸型上均无碰撞且可制造,比人工设计快10倍以上
- 支持说话与倾听双重行为,适合大规模部署的社交机器人
拟人化面部是社交交互机器人的核心,通过面部运动实现丰富的非语言交流。然而,现有拟人化面部多为定制系统:每个新脸型需大量手动机械重设计,导致大规模个性化成本高、耗时长。本文提出自动化、可扩展的机械面部生成方法,旨在快速为多样脸型生成物理可行的内部机构。我们设计了一种参数化、连杆驱动的机械面部模板,其拓扑结构与执行器布局显式参数化,支持跨不同面部形态的系统性缩放与重定向。在此基础上,提出一种分层自动设计算法:输入单张2D肖像,重建目标3D人脸,并合成无碰撞、可制造的内部机构。算法融合解剖引导的可行运动空间、基于动作单元(AU)的轨迹表达目标,以及基于碰撞的外层优化策略。此外,我们认为未来大规模部署的机械面部必须支持双向、多轮对话,而不仅是发声或倾听头。为此,开发了双身份对话式面部运动合成框架,联合建模音频驱动的说话与倾听行为,生成适合物理执行的时间一致3D面部运动。通过大量实验验证系统性能,包括(i)在多种脸型上对自动机制合成的定量评估,(ii)与人工设计的对比,(iii)对话式面部运动合成基准测试与实时部署,(iv)感知用户研究。
原文摘要 · Abstract (English)
Animatronic faces are a central component of socially interactive robots, enabling rich nonverbal communication through facial articulation. However, state-of-the-art animatronic faces are typically tailored systems: each new facial geometry requires extensive manual mechanical redesign, making large-scale personalization prohibitively slow and costly. In this work, we pursue automated and scalable mechanical face synthesis, aiming to rapidly generate a physically realizable facial mechanism for a wide range of facial geometries. We introduce a parametric, linkage-driven mechanical face template whose topology and actuator layout are explicitly parameterized to support systematic scaling and retargeting across diverse facial morphologies. Building on this template, we propose a hierarchical automatic design algorithm that takes a single 2D portrait as input, reconstructs a target 3D face, and synthesizes a collision-free, manufacturable internal mechanism. The algorithm combines anatomy-guided feasible motion volumes, Action Unit (AU)-derived trajectory-based expressiveness objectives, and a collision-driven outer-loop refinement strategy. Beyond hardware synthesis, we argue that future mechanical faces deployed at scale must engage in bidirectional, multi-turn conversation rather than functioning solely as speaking or listening heads. To this end, we develop a dual-identity conversational facial motion synthesis framework that jointly models speaking and listening behaviors from audio, producing temporally coherent 3D facial motion suitable for physical execution. We validate our system through extensive experiments, including (i) quantitative evaluation of automatic mechanism synthesis across diverse facial geometries, (ii) comparisons against manual mechanical design, (iii) benchmarks on conversational facial motion synthesis and real-time deployment, and (iv) perceptual user studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。