让手语视频带情绪,提升自然度与表现力
EASL: Multi-Emotion Guided Semantic Disentanglement for Expressive Sign Language Generation
- 分离语义与情感特征,分步训练实现精准控制
- 生成7类情绪的签语动作,准确率优于所有基线
- 适配扩散模型,适合需要高表达力的手语生成
大型语言模型已推动手语生成技术发展,可将文本自动转化为高质量手语视频,为聋人群体提供无障碍沟通。然而现有基于LLM的方法侧重语义准确性,忽略情感表达,导致输出缺乏自然性和表现力。本文提出EASL(情感感知手语生成)架构,通过多情绪引导的语义解耦机制,实现细粒度情感融合。引入渐进式训练的情感-语义解耦模块,分别提取语义与情感特征。在姿态解码阶段,情感表征引导语义交互,生成具有7类情绪置信度评分的手语动作,支持情感识别。实验表明,EASL在姿态准确率上全面超越对比基线,有效适配扩散模型,生成更具表现力的手语视频。
原文摘要 · Abstract (English)
Large language models have revolutionized sign language generation by automatically transforming text into high-quality sign language videos, providing accessible communication for the Deaf community. However, existing LLM-based approaches prioritize semantic accuracy while overlooking emotional expressions, resulting in outputs that lack naturalness and expressiveness. We propose EASL (Emotion-Aware Sign Language), a multi-emotion-guided generation architecture for fine-grained emotional integration. We introduce emotion-semantic disentanglement modules with progressive training to separately extract semantic and affective features. During pose decoding, the emotional representations guide semantic interaction to generate sign poses with 7-class emotion confidence scores, enabling emotional expression recognition. Experimental results demonstrate that EASL achieves pose accuracy superior to all compared baselines by integrating multi-emotion information and effectively adapts to diffusion models to generate expressive sign language videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。