arXiv:2512.04048cs.CVcs.CL2025-12ACL被引 3

提出Stable Signer模型,端到端生成高质量手语视频。

Stable Signer: Hierarchical Sign Language Generative Model

  • 将手语生成重构为文本理解与动作生成的分层端到端任务
  • 引入SLUL和SLP-MoE模块,生成效果比当前最优方法提升48.6%
  • 适合需要高保真多风格手语视频生成的研究与应用

手语生成(SLP)是将复杂文本转化为真实视频的过程。以往研究多聚焦于文本到词汇(Text2Gloss)、词汇到姿态(Gloss2Pose)、姿态到视频(Pose2Vid)等阶段,部分工作关注提示到词汇(Prompt2Gloss)和文本到虚拟人(Text2Avatar)阶段。然而,由于文本转换不准确、姿态生成偏差及姿态渲染为真人视频的失真问题,各阶段误差累积导致该领域进展缓慢。为此,本文简化传统冗余结构,优化任务目标,提出新的手语生成模型Stable Signer。该模型将SLP任务重构为仅包含文本理解(Prompt2Gloss、Text2Gloss)与Pose2Vid的分层端到端生成流程,通过提出的手语理解链接器SLUL执行文本理解,并利用命名为SLP-MoE的手势渲染专家模块生成手势,实现高质量、多风格手语视频的端到端生成。SLUL采用新设计的语义感知词汇掩码损失(SAGM Loss)进行训练,性能相比当前最先进方法提升48.6%。

原文摘要 · Abstract (English)

Sign Language Production (SLP) is the process of converting the complex input text into a real video. Most previous works focused on the Text2Gloss, Gloss2Pose, Pose2Vid stages, and some concentrated on Prompt2Gloss and Text2Avatar stages. However, this field has made slow progress due to the inaccuracy of text conversion, pose generation, and the rendering of poses into real human videos in these stages, resulting in gradually accumulating errors. Therefore, in this paper, we streamline the traditional redundant structure, simplify and optimize the task objective, and design a new sign language generative model called Stable Signer. It redefines the SLP task as a hierarchical generation end-to-end task that only includes text understanding (Prompt2Gloss, Text2Gloss) and Pose2Vid, and executes text understanding through our proposed new Sign Language Understanding Linker called SLUL, and generates hand gestures through the named SLP-MoE hand gesture rendering expert block to end-to-end generate high-quality and multi-style sign language videos. SLUL is trained using the newly developed Semantic-Aware Gloss Masking Loss (SAGM Loss). Its performance has improved by 48.6% compared to the current SOTA generation methods.

手语生成端到端多风格姿态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。