arXiv:2603.10446cs.CV2026-03中稿 · ECCV被引 1

用稀疏关键帧生成流畅多语种手语,解决动作模糊与不连贯问题。

SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning

  • 通过稀疏关键帧捕捉手语运动规律,从离散锚点生成连续动作。
  • 在4种手语上实现最优性能,关键帧可精准编辑时空轨迹。
  • 支持多语种扩展与真实感渲染,适合无障碍沟通与内容创作。

手语生成面临核心矛盾:直接文本到动作模型易出现回归均值现象,而字典检索方法则导致动作断续。为此,我们提出一种新型训练范式,利用稀疏关键帧捕获人类手语的潜在运动分布。通过从离散锚点生成密集动作,该方法缓解了回归均值问题并确保动作连贯性。为实现规模化,我们引入FAST——一个超高效的手语分割模型,可自动挖掘精确的时间边界。随后提出SignSparK,一种基于条件流匹配(CFM)的框架,利用这些时间锚点合成3D手语序列。该关键帧驱动范式还实现了关键帧到动作(KF2P)生成,使手语序列的时空编辑成为可能。此外,SignSparK可跨4种不同手语语言,构成目前规模最大的多语言手语生成框架,并集成3D高斯泼溅实现逼真渲染。大量实验表明,SignSparK在多种手语生成任务和多语言基准上均达到最先进水平。

原文摘要 · Abstract (English)

Sign Language Production (SLP) faces a fundamental trade-off: direct text-to-pose models suffer from regression-to-the-mean effects, while dictionary-retrieval methods produce disjointed transitions. To resolve this, we propose a novel training paradigm that leverages sparse keyframes to capture the underlying kinematic distribution of human signing. By generating dense motion from discrete anchors, our approach mitigates regression-to-the-mean while ensuring fluid articulation. To achieve this at scale, we introduce FAST, an ultra-efficient sign segmentation model that automatically mines precise temporal boundaries. We then present SignSparK, a Conditional Flow Matching (CFM) framework that utilizes these temporal anchors to synthesize 3D signing sequences. This keyframe-driven formulation also unlocks Keyframe-to-Pose (KF2P) generation, making precise spatiotemporal editing of signing sequences possible. Furthermore, SignSparK scales across four distinct sign languages, constituting the largest multilingual SLP framework to date, and integrates 3D Gaussian Splatting for photorealistic rendering. Extensive evaluations demonstrate that SignSparK achieves state-of-the-art across diverse SLP tasks and multilingual benchmarks. Our code is available at https://github.com/JianHe0628/SignSparK.

手语生成多语言关键帧3D渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。