arXiv:2602.07506cs.ROcs.AI2026-02中稿 · the 2026 IEEE Inte…被引 1

让机器人实时逼真模仿人脸表情,0.05秒内完成

VividFace: Real-Time and Realistic Facial Expression Shadowing for Humanoid Robots

  • 优化的X2CNet++框架提升表情迁移细节
  • 0.05秒内完成实时表情同步,支持多脸型适配
  • 适合人机交互、仿生机器人等场景

类人机器人面部表情模仿可实现实时、逼真的面部表情复现,对提升人机交互自然性至关重要。现有方法受限于离线视频推理设计,难以兼顾实时性与表现力,且对细微表情捕捉能力不足。为此,本文提出VividFace系统,通过优化的仿生表情迁移框架X2CNet++,结合特征自适应训练策略,实现跨源图像间更精准的表情对齐。同时,采用视频流兼容的推理管道与异步输入输出机制,保障高效设备通信。系统可在0.05秒内生成生动逼真的类人表情,并在多种面部配置下保持泛化能力。大量真实场景演示验证了其实用性。视频展示见:https://lipzh5.github.io/VividFace/

原文摘要 · Abstract (English)

Humanoid facial expression shadowing enables robots to realistically imitate human facial expressions in real time, which is critical for lifelike, facially expressive humanoid robots and affective human-robot interaction. Existing progress in humanoid facial expression imitation remains limited, often failing to achieve either real-time performance or realistic expressiveness due to offline video-based inference designs and insufficient ability to capture and transfer subtle expression details. To address these limitations, we present VividFace, a real-time and realistic facial expression shadowing system for humanoid robots. An optimized imitation framework X2CNet++ enhances expressiveness by fine-tuning the human-to-humanoid facial motion transfer module and introducing a feature-adaptation training strategy for better alignment across different image sources. Real-time shadowing is further enabled by a video-stream-compatible inference pipeline and a streamlined workflow based on asynchronous I/O for efficient communication across devices. VividFace produces vivid humanoid faces by mimicking human facial expressions within 0.05 seconds, while generalizing across diverse facial configurations. Extensive real-world demonstrations validate its practical utility. Videos are available at: https://lipzh5.github.io/VividFace/.

人脸模仿机器人交互实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。