arXiv:2508.14574cs.CLcs.LG2025-08

用四元数和对比学习减少手语生成中的姿势噪声,提升准确性

Towards Skeletal and Signer Noise Reduction in Sign Language Production via Quaternion-Based Pose Encoding and Contrastive Learning

  • 用四元数编码骨骼旋转,优化关节角度精度
  • 对比损失使模型忽略无关的形态与风格差异,正确关键点率提升16%
  • 适合关注手语生成鲁棒性与语义对齐的研究者

神经手语生成中的主要挑战是手势在类内存在显著差异,源于签名者体型和训练数据中的风格多样性。为提升对这些差异的鲁棒性,本文对标准渐进式变压器(PT)架构提出两项改进:首先,采用四元数空间中的骨旋转编码姿势,并使用测地线损失训练,以提高关节运动的角度准确性和清晰度;其次,引入对比损失,通过词汇重叠或SBERT句子相似度构建解码器嵌入的语义结构,旨在过滤掉不传递语义信息的解剖特征与风格特征。在Phoenix14T数据集上,仅使用对比损失即带来16%的正确关键点概率提升;结合四元数编码后,平均骨角误差降低6%。结果表明,在基于Transformer的手语生成模型中融入骨骼结构建模与语义引导的对比目标具有显著优势。

原文摘要 · Abstract (English)

One of the main challenges in neural sign language production (SLP) lies in the high intra-class variability of signs, arising from signer morphology and stylistic variety in the training data. To improve robustness to such variations, we propose two enhancements to the standard Progressive Transformers (PT) architecture (Saunders et al., 2020). First, we encode poses using bone rotations in quaternion space and train with a geodesic loss to improve the accuracy and clarity of angular joint movements. Second, we introduce a contrastive loss to structure decoder embeddings by semantic similarity, using either gloss overlap or SBERT-based sentence similarity, aiming to filter out anatomical and stylistic features that do not convey relevant semantic information. On the Phoenix14T dataset, the contrastive loss alone yields a 16% improvement in Probability of Correct Keypoint over the PT baseline. When combined with quaternion-based pose encoding, the model achieves a 6% reduction in Mean Bone Angle Error. These results point to the benefit of incorporating skeletal structure modeling and semantically guided contrastive objectives on sign pose representations into the training of Transformer-based SLP models.

手语生成四元数对比学习姿态编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。