让手语生成更自然,通过几何约束提升身体动作真实度
Geometry-Aware Losses for Structure-Preserving Text-to-Sign Language Generation
- 引入骨骼关节几何约束,建模肩臂手的协调关系
- 相比最佳基线,动作真实度提升56.51%,骨长与运动差异下降超18%
- 适合手语生成、具身智能等需要自然人体动作的场景
从文本到视频的手语翻译对聋哑人士沟通至关重要。核心挑战在于生成准确且自然的身体姿态与动作以忠实传达语义。现有方法常忽略人体骨骼运动的解剖约束与协同模式,导致动作僵硬或生物力学上不合理。为此,我们提出一种新方法,显式建模肩、臂、手等骨骼关节间的关系,通过施加关节位置、骨长及运动动态的几何约束。训练中引入父节点相对重加权机制,增强手指灵活性并降低动作僵硬度。同时,骨姿态损失与骨长约束确保解剖结构一致性。实验表明,该方法将先前最优结果与真值基准之间的性能差距缩小了56.51%,骨长差异与运动方差分别降低18.76%和5.48%,显著提升了动作的真实感与自然度。
原文摘要 · Abstract (English)
Sign language translation from text to video plays a crucial role in enabling effective communication for Deaf and hard--of--hearing individuals. A major challenge lies in generating accurate and natural body poses and movements that faithfully convey intended meanings. Prior methods often neglect the anatomical constraints and coordination patterns of human skeletal motion, resulting in rigid or biomechanically implausible outputs. To address this, we propose a novel approach that explicitly models the relationships among skeletal joints--including shoulders, arms, and hands--by incorporating geometric constraints on joint positions, bone lengths, and movement dynamics. During training, we introduce a parent-relative reweighting mechanism to enhance finger flexibility and reduce motion stiffness. Additionally, bone-pose losses and bone-length constraints enforce anatomically consistent structures. Our method narrows the performance gap between the previous best and the ground-truth oracle by 56.51%, and further reduces discrepancies in bone length and movement variance by 18.76% and 5.48%, respectively, demonstrating significant gains in anatomical realism and motion naturalness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。