arXiv:2608.13368cs.CVcs.AI2026-08

用多专家GAN生成手语视频,提升听障人士沟通体验。

Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs

论文配图:Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs
图 1 · 摘自论文原文
  • 设计三个专用判别器分别指导生成器的全局、手部和头部区域。
  • 1.3B参数模型达30.7 PSNR、0.965 SSIM,可在消费级显卡运行。
  • 采用自适应特征融合与动态损失机制,提升生成稳定性和细节。

本文提出一种基于损失引导多专家GAN的手语视频合成框架,旨在改善听障人士的沟通体验。通过全局、手部和头部三个专用判别器,分别引导生成器对应区域的特征学习,实现隐式特征专业化,无需额外多样性损失。为稳定多判别器系统早期训练的混沌动态,引入权重为10%的联合损失一致性机制,使各判别器向整体平均值对齐。生成器采用双路径卷积-变压器结构,结合可学习的自适应特征融合模块,兼顾卷积的稳定性与窗口自注意力的细节表现力。训练采用交替三模式策略(判别器、整体生成、分支特化生成)。在自建156GB数据集上,0.2B参数版本达29.8 PSNR(0.959 SSIM),1.3B版本达30.7 PSNR(0.965 SSIM),推理显存占用分别为1.5 GB和8 GB,支持消费级硬件部署。完整消融实验因单卡训练需2-3个月仍在进行中。系统已在2025年香港前沿科技峰会上展示。

原文摘要 · Abstract (English)

This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialized discriminators -- global, hand, and head -- each guide a corresponding expert branch in the generator toward a distinct visual region, enabling implicit feature specialization without explicit diversity losses. To stabilize this multi-discriminator system, whose early-phase training otherwise exhibits chaotic dynamics, we introduce a United Loss consensus mechanism that regularizes each discriminator toward the ensemble average at a 10% weight. Each branch further adopts a dual-pathway convolutional-transformer design with learnable AdaptiveFeatureFusion, balancing the stability of convolutions against the detail of windowed self-attention. The generator is trained using an alternating three-mode schedule (discriminator, holistic generation, branch-specialized generation). On a custom 156GB dataset with a filtered test set that removes easy and repetitive samples, our 0.2B-parameter variant achieves 29.8 PSNR (0.959 SSIM) and the 1.3B-parameter variant achieves 30.7 PSNR (0.965 SSIM), with inference VRAM footprints of 1.5 GB and 8 GB respectively, enabling deployment on consumer-grade hardware. Full ablation studies remain ongoing due to the 2-3 month training cycle on a single GPU. The system was showcased at the 2025 Hong Kong Frontier Technology Summit.

手语生成GAN视频合成多专家

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。