arXiv:2506.11621cs.CV2025-06被引 5

用多模态对齐生成更自然的手语视频

SignAligner: Harmonizing Complementary Pose Modalities for Coherent Sign Language Generation

  • 联合生成姿势、手势和身体动作,融合文本语义
  • 在线协同修正消除时空冲突,提升一致性
  • 适用于手语合成、无障碍交流系统研发

手语生成旨在根据口语生成多样化的手语表达。由于手语包含复杂的手势、面部表情和身体动作,实现真实自然的生成仍具挑战。本文引入PHOENIX14T+数据集,扩展自RWTH-PHOENIX-Weather 2014T,新增三种手语表示:Pose、Hamer和Smplerx。提出SignAligner方法,分三阶段:基于文本驱动的多模态联合生成,利用Transformer编码器提取语义特征,通过跨模态注意力生成多样化表示;在线协同修正阶段采用动态损失加权与跨模态注意力,增强模态间互补性,消除时空冲突,确保语义连贯与动作一致;最后将修正后的姿态输入预训练视频生成网络,生成高保真手语视频。大量实验表明,SignAligner显著提升生成视频的准确性和表现力。

原文摘要 · Abstract (English)

Sign language generation aims to produce diverse sign representations based on spoken language. However, achieving realistic and naturalistic generation remains a significant challenge due to the complexity of sign language, which encompasses intricate hand gestures, facial expressions, and body movements. In this work, we introduce PHOENIX14T+, an extended version of the widely-used RWTH-PHOENIX-Weather 2014T dataset, featuring three new sign representations: Pose, Hamer and Smplerx. We also propose a novel method, SignAligner, for realistic sign language generation, consisting of three stages: text-driven pose modalities co-generation, online collaborative correction of multimodality, and realistic sign video synthesis. First, by incorporating text semantics, we design a joint sign language generator to simultaneously produce posture coordinates, gesture actions, and body movements. The text encoder, based on a Transformer architecture, extracts semantic features, while a cross-modal attention mechanism integrates these features to generate diverse sign language representations, ensuring accurate mapping and controlling the diversity of modal features. Next, online collaborative correction is introduced to refine the generated pose modalities using a dynamic loss weighting strategy and cross-modal attention, facilitating the complementarity of information across modalities, eliminating spatiotemporal conflicts, and ensuring semantic coherence and action consistency. Finally, the corrected pose modalities are fed into a pre-trained video generation network to produce high-fidelity sign language videos. Extensive experiments demonstrate that SignAligner significantly improves both the accuracy and expressiveness of the generated sign videos.

手语生成多模态对齐视频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。