arXiv:2412.03878cs.CV2024-12ECCV被引 4

用AI生成逼真可定制的手语视频,让聋哑人群也能看懂全球影视内容。

DiffSign: AI-Assisted Generation of Customizable Sign Language Videos With Enhanced Realism

  • 结合参数化与生成模型,用3D动作重定向生成高保真手语姿态。
  • 生成视频在时间一致性与真实感上优于仅用文本提示的扩散模型。
  • 支持图像提示和多模态输入,可自定义手语者外貌特征,适合无障碍传播。

近年来,流媒体服务的普及使全球观众能观看相同的影视内容。尽管翻译与配音服务逐渐增加,但针对听障及重听(DHH)群体的内容可访问性仍滞后。本文旨在通过生成逼真且富有表现力的合成手语者,提升媒体内容对DHH群体的可及性。为避免单一手语者在全球使用中缺乏吸引力,我们结合参数建模与生成建模,实现外观可定制的合成手语者生成。首先,通过优化参数模型将真人手语动作重定向至3D手语虚拟人;再利用渲染出的高保真动作作为条件,驱动基于扩散模型生成的合成手语者动作。合成手语者的外观由图像提示通过视觉适配器控制。实验表明,本方法生成的手语视频在时间一致性和真实感方面优于仅依赖文本提示的扩散模型。同时支持多模态提示,允许用户根据肤色、性别等多样性特征进一步定制手语者形象。该方法还可用于手语者匿名化处理。

原文摘要 · Abstract (English)

The proliferation of several streaming services in recent years has now made it possible for a diverse audience across the world to view the same media content, such as movies or TV shows. While translation and dubbing services are being added to make content accessible to the local audience, the support for making content accessible to people with different abilities, such as the Deaf and Hard of Hearing (DHH) community, is still lagging. Our goal is to make media content more accessible to the DHH community by generating sign language videos with synthetic signers that are realistic and expressive. Using the same signer for a given media content that is viewed globally may have limited appeal. Hence, our approach combines parametric modeling and generative modeling to generate realistic-looking synthetic signers and customize their appearance based on user preferences. We first retarget human sign language poses to 3D sign language avatars by optimizing a parametric model. The high-fidelity poses from the rendered avatars are then used to condition the poses of synthetic signers generated using a diffusion-based generative model. The appearance of the synthetic signer is controlled by an image prompt supplied through a visual adapter. Our results show that the sign language videos generated using our approach have better temporal consistency and realism than signing videos generated by a diffusion model conditioned only on text prompts. We also support multimodal prompts to allow users to further customize the appearance of the signer to accommodate diversity (e.g. skin tone, gender). Our approach is also useful for signer anonymization.

手语生成扩散模型虚拟形象无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。