用2D关键点生成手语视频,显著提升清晰度与可懂性
SignRefine: Adapting Foundational Video Models for Sign Language Generation

- 基于预训练视频扩散模型,通过局部适配器精修手部和面部区域
- 在手姿精度上比最强基线提升30%,80%以上用户认为视觉质量更优
- 适用于手语生成、无障碍交流,尤其适合残障辅助技术研究者
手语视频生成需要精确的手部和面部动作表达,但当前主要以口语视频训练的视频扩散模型会产生影响理解的伪影。我们提出SignRefine,仅依赖2D关键点条件即可生成可理解的手语视频,且能适应不同外观和视觉环境。该方法基于预训练视频扩散变换器,引入带有空间定位的局部适配器,有选择性地精修手部与面部区域,引导强基模型先验朝向准确表达。为支持此项工作及更广泛的手语研究,我们构建了NVSign——一个大规模原生手语视频数据集,涵盖多样签名者外观、环境与自然对话场景。在该数据集上训练后,模型在手姿精度指标上相较最强基线最高提升30%,并在超过80%的对比中被手语使用者评为视觉质量和可懂性更优。
原文摘要 · Abstract (English)
Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models, trained predominantly on spoken-language video, produce artifacts that render signing unintelligible. We propose SignRefine, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions. Our approach builds on a pretrained video diffusion transformer and introduces local adapters with spatial grounding to selectively refine hand and face regions, steering the strong base model's prior toward accurate articulation. To enable this work and support broader sign language research, we present NVSign, a large-scale dataset of video content natively produced in sign language, offering diverse signer appearances, environments, and natural conversational settings. Trained on this data, our model shows up to 30% improvement in hand pose precision metrics over the strongest baseline and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。