arXiv:2601.20791cs.CVcs.AI2026-01被引 1

无需微调,通过几何变换消除文本到视频生成中的性别偏见

FAIRT2V: Training-Free Debiasing for Text-to-Video Diffusion Models

  • 用锚点球面变换中和提示词嵌入,消除预训练编码器的隐含性别关联
  • 在早期生成阶段动态去偏,使视频质量下降不足2%,职业偏见降低超60%
  • 适合关注AI公平性的研究人员与内容生成平台开发者

文本到视频(T2V)扩散模型发展迅速,但其种族与性别偏见尚未得到充分研究。本文提出FairT2V,一种无需微调的训练自由去偏框架,可有效缓解由预训练文本编码器引发的偏见。我们分析发现,该偏见主要源于编码器对中性提示也编码隐含性别关联。通过构建性别倾向评分量化此现象,并基于此设计锚点引导的球面测地线变换,对提示嵌入进行中和处理以保留语义。为保障时间连贯性,仅在生成初期的身份形成阶段应用去偏策略,结合动态去噪调度。进一步提出融合VideoLLM推理与人工验证的视频级公平性评估协议。在现代T2V模型Open-Sora上的实验表明,FairT2V显著降低各职业场景下的性别偏见,同时对视频质量影响极小。

原文摘要 · Abstract (English)

Text-to-video (T2V) diffusion models have achieved rapid progress, yet their demographic biases, particularly gender bias, remain largely unexplored. We present FairT2V, a training-free debiasing framework for text-to-video generation that mitigates encoder-induced bias without finetuning. We first analyze demographic bias in T2V models and show that it primarily originates from pretrained text encoders, which encode implicit gender associations even for neutral prompts. We quantify this effect with a gender-leaning score that correlates with bias in generated videos. Based on this insight, FairT2V mitigates demographic bias by neutralizing prompt embeddings via anchor-based spherical geodesic transformations while preserving semantics. To maintain temporal coherence, we apply debiasing only during early identity-forming steps through a dynamic denoising schedule. We further propose a video-level fairness evaluation protocol combining VideoLLM-based reasoning with human verification. Experiments on the modern T2V model Open-Sora show that FairT2V substantially reduces demographic bias across occupations with minimal impact on video quality.

文本生成视频去偏扩散模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。