arXiv:2607.14836cs.CV2026-07

用物理约束提升手语生成的自然度与准确性

Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation

论文配图:Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation
图 1 · 摘自论文原文
  • 引入可微分几何模块,实时校正骨骼长度和关节角度
  • 在PHOENIX14T和CSL-Daily上关键指标优于基线模型
  • 适合需要高真实感手语动画的无障碍应用

手语生成需同时保证语义准确性和生物力学合理性。现有方法将骨架视为无结构向量,导致骨长漂移、关节角度异常及手指短暂锁定。本文提出PIDiffSign,一种融合解剖约束的扩散模型,采用Transformer编码器-解码器架构,解码器通过自适应零初始化层归一化和跨注意力机制条件于扩散时间步。引入可微分几何模块,在生成过程中持续强制骨骼长度一致性和生物学合理的关节角度。训练结合人体解剖、运动学、角度及手指关节约束,并使用对比式词-姿态对齐损失与无分类器引导实现语义条件采样。在PHOENIX14T和CSL-Daily数据集上的实验表明,该模型在姿态精度、关节角度正确性、分布真实性及反向翻译质量上均显著优于强基线扩散模型,验证了物理信息扩散能有效提升手语生成的运动真实感与语义保真度。

原文摘要 · Abstract (English)

Sign language production, which generates continuous 3D skeletal motion from spoken language input, must simultaneously satisfy two constraints: semantic fidelity, so that a deaf viewer can recognize the intended sequence of glosses, and biomechanical plausibility, so that the generated skeleton respects anatomical constraints. Existing approaches optimize semantic reconstruction through coordinate-based objectives that treat the skeleton as an unstructured vector, thus allowing for bone length drift, joint angle violations, and temporarily locked fingers. We introduce PIDiffSign, a physics-informed diffusion model for gloss-to-pose translation that incorporates anatomical constraints into both the architecture and training objective. The model uses a Transformer encoder-decoder, where the decoder is conditioned on the diffusion time step through adaptive zero-initialized layer normalization and cross-attends to gloss representations. A differentiable geometry module enforces bone length consistency and biologically valid joint angles throughout generation. Training combines anthropomorphic, kinematic, angular, and finger-joint constraints with a contrastive gloss-pose alignment loss and classifier-free guidance for semantically conditioned sampling. Experiments on the PHOENIX14T and CSL-Daily benchmarks show consistent improvements over a strong diffusion baseline in pose accuracy, joint-angle correctness, distributional realism, and back-translation quality. These results demonstrate that physics-informed diffusion improves both motion realism and semantic fidelity for sign language generation.

手语生成扩散模型生物力学约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。