用扩散模型生成手语视频的静态插图,助力教育可视化。
IlluSign: Illustrating Sign Language Videos by Leveraging the Attention Mechanism
- 通过干预扩散模型注意力层,融合手势与表情的几何和语义信息。
- 将手语起始帧与结束帧合成一张插图,用箭头标注动作方向。
- 自动化生成高质量手语插图,降低人工绘图成本,适合教学使用。
手语是包含手势与面部表情等非手动元素的动态视觉语言。尽管手语视频常用于教学与记录,其动态特性使学习者与教师难以深入分析。本文提出一种方法,将手语视频转换为静态插图,作为视频内容的补充教育资源。传统上该过程依赖艺术家手工绘制,成本高昂。我们利用生成模型理解图像的语义与几何特征,实现手语视频到素描风格插图的转化。方法将一个手语动作的起始帧与结束帧合并为一张插图,并以箭头表示手部运动方向。针对手语中手势与表情的复杂性,我们在扩散模型的去噪过程中干预高分辨率注意力层,将风格作为键值注入,几何与边缘信息作为查询融合。最终插图通过整合起始与结束帧的注意力权重,实现平滑融合。本方法在推理时可低成本生成手语插图,填补教学资源空白。
原文摘要 · Abstract (English)
Sign languages are dynamic visual languages that involve hand gestures, in combination with non manual elements such as facial expressions. While video recordings of sign language are commonly used for education and documentation, the dynamic nature of signs can make it challenging to study them in detail, especially for new learners and educators. This work aims to convert sign language video footage into static illustrations, which serve as an additional educational resource to complement video content. This process is usually done by an artist, and is therefore quite costly. We propose a method that illustrates sign language videos by leveraging generative models' ability to understand both the semantic and geometric aspects of images. Our approach focuses on transferring a sketch like illustration style to video footage of sign language, combining the start and end frames of a sign into a single illustration, and using arrows to highlight the hand's direction and motion. While many style transfer methods address domain adaptation at varying levels of abstraction, applying a sketch like style to sign languages, especially for hand gestures and facial expressions, poses a significant challenge. To tackle this, we intervene in the denoising process of a diffusion model, injecting style as keys and values into high resolution attention layers, and fusing geometric information from the image and edges as queries. For the final illustration, we use the attention mechanism to combine the attention weights from both the start and end illustrations, resulting in a soft combination. Our method offers a cost effective solution for generating sign language illustrations at inference time, addressing the lack of such resources in educational materials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。