arXiv:2508.04049cs.CV2025-08被引 2

让手势动作独立于具体人,实现任意演员的自然手语视频生成

Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation

  • 用一个动作片段构建无身份绑定的手势词典,降低数据需求
  • 将词典动作转为连贯运动轨迹,再渲染成任意真人形象的视频
  • 适合需要灵活定制手语角色的研究与应用

手语视频生成需在精确语义控制下产出自然动作与真实外观,但面临两个核心挑战:依赖大量特定表演者数据,且泛化能力差。本文提出一种新范式,通过两阶段合成框架解耦动作语义与表演者身份。首先,构建无需身份信息的多模态动作词典,每个手势仅需一次录制即可存储为姿态、手势与3D网格序列;其次,设计离散到连续的动作生成阶段,将检索出的手势序列转化为时间一致的运动轨迹,并通过身份感知神经渲染生成任意表演者的逼真视频。相比以往受限于特定数据集的方法,本方法将动作视为第一性元素:学习到的潜在姿态动态可作为可移植的“编舞层”,适配不同外貌的人体形象。大量实验表明,解耦动作与身份不仅可行,且显著提升生成质量与个性化灵活性。

原文摘要 · Abstract (English)

Sign language video generation requires producing natural signing motions with realistic appearances under precise semantic control, yet faces two critical challenges: excessive signer-specific data requirements and poor generalization. We propose a new paradigm for sign language video generation that decouples motion semantics from signer identity through a two-phase synthesis framework. First, we construct a signer-independent multimodal motion lexicon, where each gloss is stored as identity-agnostic pose, gesture, and 3D mesh sequences, requiring only one recording per sign. This compact representation enables our second key innovation: a discrete-to-continuous motion synthesis stage that transforms retrieved gloss sequences into temporally coherent motion trajectories, followed by identity-aware neural rendering to produce photorealistic videos of arbitrary signers. Unlike prior work constrained by signer-specific datasets, our method treats motion as a first-class citizen: the learned latent pose dynamics serve as a portable "choreography layer" that can be visually realized through different human appearances. Extensive experiments demonstrate that disentangling motion from identity is not just viable but advantageous - enabling both high-quality synthesis and unprecedented flexibility in signer personalization.

手语生成动作解耦神经渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。