用真实手势片段+模型精修,生成更自然连贯的手语视频。
SignRR: Retrieve and Refine Real Motion for Sign Language Production

- 先检索真实手语片段,再用模型全局优化序列
- 在PHOENIX14T上比现有方法提升1.8%的翻译准确率
- 适合需要高精度手语生成的研究者和开发者
手语生成(SLP)旨在从口语中生成连续的手语动作,通常通过词素到姿态的转换实现。以往方法主要分为两类:生成式模型从学习到的先验或噪声中合成动作,难以保留罕见手部姿态和签者特有表达;基于检索的方法复用真实、清晰的动作片段,但不同签者和共音情境下的片段拼接会导致节奏与风格不一致,不仅在边界处出现断裂。为此,我们提出「检索-精修」范式,即以真实检索到的动作为基础,通过模型精修实现全局连贯性,而非从零生成。我们的框架SignRR从真实手语片段字典初始化动作,并使用一种关注局部特征的残差量化变分自编码器(Residual VQ-VAE)对整段序列进行精修,其中残差量化保留精细手部动作,时序长度差异在隐空间中处理。在PHOENIX14T和CSL-Daily数据集上的实验表明,SignRR在回译任务中达到最新性能,同时保持优异的姿态质量。
原文摘要 · Abstract (English)
Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms. Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, making rare hand configurations and signer-specific articulation difficult to preserve. Retrieval-based methods reuse real, well-articulated motion segments, but concatenating segments from different signers and co-articulation contexts can introduce rhythm and style inconsistencies across the full sequence, not only at segment boundaries. These limitations suggest a complementary solution: use retrieval to provide realistic articulation, and use learned refinement to impose the global coherence that retrieval alone lacks. We therefore propose retrieve-and-refine, a paradigm that starts from real retrieved motion and refines it into a globally coherent signing sequence rather than generating motion from scratch. Our framework, SignRR, initializes motion from a dictionary of real sign segments and refines the full sequence with a part-aware Residual VQ-VAE, where residual quantization preserves fine hand articulation and temporal length differences are handled in the latent space. Experiments on PHOENIX14T and CSL-Daily show that SignRR achieves state-of-the-art back-translation performance while maintaining competitive pose quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。