arXiv:2504.06610cs.LGcs.CV2025-04被引 8

无需词典直接生成手语动作,通过解耦与正则化提升可解释性。

Disentangle and Regularize: Sign Language Production with Articulator-Based Disentanglement and Channel-Aware Regularization

  • 按面部、双手和躯干分块解耦手语动作特征,增强表示可读性。
  • 在PHOENIX14T和CSL-Daily上达到当前最优性能,无需词典监督。
  • 基于通道感知正则化,让模型关注不同身体部位的重要性差异。

本文提出DARSLP,一种无需词典的Transformer型手语生成框架,直接将口语文本映射为手语动作序列。首先训练一个姿态自编码器,采用基于动作器的解耦策略,将面部、右手、左手和躯干特征分别建模,以促进结构化且可解释的表征学习。随后,使用非自回归Transformer解码器从输入句子的词级文本嵌入中预测这些潜在表示。为引导训练过程,引入通道感知正则化:通过KL散度损失将预测的潜在分布与真实编码提取的先验对齐,并根据各通道对应的动作器区域加权损失贡献,使模型在训练中考虑不同动作器的相对重要性。该方法不依赖词典标注或预训练模型,在PHOENIX14T和CSL-Daily数据集上取得当前最佳结果。

原文摘要 · Abstract (English)

In this work, we propose DARSLP, a simple gloss-free, transformer-based sign language production (SLP) framework that directly maps spoken-language text to sign pose sequences. We first train a pose autoencoder that encodes sign poses into a compact latent space using an articulator-based disentanglement strategy, where features corresponding to the face, right hand, left hand, and body are modeled separately to promote structured and interpretable representation learning. Next, a non-autoregressive transformer decoder is trained to predict these latent representations from word-level text embeddings of the input sentence. To guide this process, we apply channel-aware regularization by aligning predicted latent distributions with priors extracted from the ground-truth encodings using a KL divergence loss. The contribution of each channel to the loss is weighted according to its associated articulator region, enabling the model to account for the relative importance of different articulators during training. Our approach does not rely on gloss supervision or pretrained models, and achieves state-of-the-art results on the PHOENIX14T and CSL-Daily datasets.

手语生成解耦表征非自回归姿势预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。