用3D动态发音体与协同发音建模,让机器人说话时嘴唇更自然逼真。
Realistic Lip Motion Generation Based on 3D Dynamic Viseme and Coarticulation Modeling for Human-Robot Interaction
- 基于发音体和协同发音机制生成连贯唇动轨迹
- 在14自由度机器人上实现高精度同步,PCC达0.92,MAJ低至0.03
- 适合需要真实非语言交互的具身智能机器人研发
自然的人机非语言交互依赖于逼真的唇部同步。本文提出一种基于3D动态发音体与协同发音建模的唇动生成框架。通过分析中文发音理论,构建了符合ARKit标准的3D动态发音体库,提供连贯的唇部先验轨迹。为解决连续语音流中的动作冲突,引入声母-韵母解耦与能量调制的协同发音机制。针对14自由度人形头部执行机构,设计高维唇动重定向策略,并通过皮尔逊相关系数(PCC)和平均绝对加速度(MAJ)进行定量消融实验验证。结果表明,该方法在效率与准确性上表现优异,具备轻量、高效、实用的特点。3D动态发音体库及实际部署视频已开源:https://github.com/yuesheng21/Phoneme-to-Lip-14DOF。
原文摘要 · Abstract (English)
Realistic lip synchronization is essential for the natural human-robot non-verbal interaction of humanoid robots. Motivated by this need, this paper presents a lip motion generation framework based on 3D dynamic viseme and coarticulation modeling. By analyzing Chinese pronunciation theory, a 3D dynamic viseme library is constructed based on the ARKit standard, which offers coherent prior trajectories of lips. To resolve motion conflicts within continuous speech streams, a coarticulation mechanism is developed by incorporating initial-final (Shengmu-Yunmu) decoupling and energy modulation. After developing a strategy to retarget high-dimensional spatial lip motion to a 14-DOF lip actuation system of a humanoid head platform, the efficiency and accuracy of the proposed architecture is experimentally validated and demonstrated with quantitative ablation experiments using the metrics of the Pearson Correlation Coefficient (PCC) and the Mean Absolute Jerk (MAJ). This research offers a lightweight, efficient, and highly practical paradigm for the speech-driven lip motion generation of humanoid robots. The 3D dynamic viseme library and real-world deployment videos are available at {https://github.com/yuesheng21/Phoneme-to-Lip-14DOF}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。