arXiv:2602.11861cs.LGcs.CL2026-02

用分布式潜变量建模手语生成,提升动作真实性和语义对齐。

A$^{2}$V-SLP: Alignment-Aware Variational Modeling for Disentangled Sign Language Production

  • 通过分布式潜变量学习每个手势部位的独立表示,避免确定性嵌入崩溃。
  • 在无词汇标注场景下实现最佳反向翻译性能与更真实的动作生成。
  • 适合研究手语生成、具身语言模型及多模态对齐的开发者和学者。

基于近期手语生成的结构解耦框架,我们提出A²V-SLP,一种对齐感知的变分建模方法,学习每个动作部位的解耦潜分布而非确定性嵌入。一个解耦的变分自编码器(VAE)编码真实手语姿态序列,提取各动作部位的均值和方差向量,作为非自回归Transformer训练的分布监督信号。给定文本嵌入后,Transformer预测潜变量的均值与对数方差,而VAE解码器在解码阶段通过随机采样重构最终的手语姿态序列。该范式通过分布式潜变量建模,在不丢失动作部位表征的前提下防止潜在空间坍缩。此外,引入词素注意力机制以增强语言输入与动作表达之间的对齐。实验表明,该方法在确定性潜变量回归基础上持续取得性能提升,在完全无词素标注设置下达到当前最优反向翻译表现,并显著改善了动作的真实性。

原文摘要 · Abstract (English)

Building upon recent structural disentanglement frameworks for sign language production, we propose A$^{2}$V-SLP, an alignment-aware variational framework that learns articulator-wise disentangled latent distributions rather than deterministic embeddings. A disentangled Variational Autoencoder (VAE) encodes ground-truth sign pose sequences and extracts articulator-specific mean and variance vectors, which are used as distributional supervision for training a non-autoregressive Transformer. Given text embeddings, the Transformer predicts both latent means and log-variances, while the VAE decoder reconstructs the final sign pose sequences through stochastic sampling at the decoding stage. This formulation maintains articulator-level representations by avoiding deterministic latent collapse through distributional latent modeling. In addition, we integrate a gloss attention mechanism to strengthen alignment between linguistic input and articulated motion. Experimental results show consistent gains over deterministic latent regression, achieving state-of-the-art back-translation performance and improved motion realism in a fully gloss-free setting.

手语生成变分模型解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。