用语义对齐的向量量化提升手语生成精度与语义一致性
SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation

- 通过句级对齐与词级监督构建语义结构化的残差向量空间
- 在375小时数据上实现最佳姿态准确率,生成动作与文本语义一致
- 适合手语生成、人机交互及无障碍沟通系统研究者
美国手语(ASL)生成因配对的文本-手语动作数据有限,且难以学习既精确重建又可从语言输入预测的动作表示而面临挑战。现有方法依赖于仅优化重建效果的动作分词器,缺乏来自配对文本的显式语义监督,导致所学的标记无法有效支持语义一致且细粒度的手语生成。为此,本文提出SeRV(语义对齐残差向量量化),一种面向手语生成的语义对齐的残差向量量化分词器。SeRV通过结合句子级动作-文本对齐与标记级文本条件监督,学习一个具有语义结构的残差标记空间。在此分词器基础上,采用分层GPT以粗到精的方式预测残差动作标记,生成结构连贯且语义对齐的3D手语动作。此外,我们通过恢复YouTube-ASL视频中的配对3D动作,构建了一个大规模重建的3D手语动作-文本基准数据集。在375小时的手语视频上的实验表明,SeRV在How2Sign和YouTube-ASL数据集上均达到最先进的姿态准确率,能够直接从文本生成语义一致的3D手语动作。
原文摘要 · Abstract (English)
American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion representations both precise for reconstruction and predictable from linguistic input. Existing methods rely on motion tokenizers optimized for reconstruction, without explicit semantic supervision from paired text. As a result, the learned tokens remain limited in supporting semantically consistent and fine-grained ASL motion generation. To address this limitation, we propose SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation. SeRV learns a semantically structured residual token space by combining sentence-level motion-text alignment with token-level text-conditioned supervision. Building on this tokenizer, a Hierarchical GPT predicts residual motion tokens in a coarse-to-fine manner, generating structurally coherent and semantically aligned 3D ASL motion. We further construct a large-scale reconstructed 3D ASL motion-text benchmark by recovering paired 3D motion from YouTube-ASL videos. Experiments across 375 hours of ASL video show that SeRV achieves state-of-the-art pose accuracy on both How2Sign and YouTube-ASL datasets, while producing semantically consistent 3D ASL motion directly from text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。