arXiv:2510.17247cs.CLcs.CV2025-10被引 1

研究发现对齐训练会放大视频生成中的社会偏见,且使偏见更持久。

From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models

  • 构建事件驱动提示框架,分离动作与人物属性,精准诊断偏见来源。
  • 发现对齐训练使种族和性别偏见增强并随时间稳定,生成更流畅但刻板的视频。
  • 适合关注公平性、伦理的AI研发者及视频生成系统评估人员。

近期视频扩散模型在文本到视频生成方面取得显著进展,尤其通过基于人类偏好训练的奖励模型进行对齐微调。尽管此类方法提升了视觉质量,却可能无意中编码并放大社会偏见。为系统追踪偏见在对齐流程中的演变,我们提出VideoBiasEval——一个全面的诊断框架,用于评估视频生成中的社会表征。该框架基于成熟的社交偏见分类体系,采用事件驱动提示策略,将语义内容(动作与情境)与演员属性(性别与种族)解耦。进一步引入多粒度指标,评估:(1) 总体种族偏见,(2) 条件于种族的性别偏见,(3) 不同模型变体间社会属性的分布变化,以及(4) 偏见在视频中的时间持续性。利用该框架,我们首次完成从人类偏好数据集中的偏见,到奖励模型中的放大,再到对齐微调后视频扩散模型中传播的全链条分析。结果表明,对齐训练不仅强化了表征偏见,还使其在时间维度上更加稳定,导致生成视频虽更平滑却更具刻板印象。这些发现凸显了在整个对齐过程中进行偏见感知评估与缓解的必要性,以保障视频生成的公平与社会责任。

原文摘要 · Abstract (English)

Recent advances in video diffusion models have significantly enhanced text-to-video generation, particularly through alignment tuning using reward models trained on human preferences. While these methods improve visual quality, they can unintentionally encode and amplify social biases. To systematically trace how such biases evolve throughout the alignment pipeline, we introduce VideoBiasEval, a comprehensive diagnostic framework for evaluating social representation in video generation. Grounded in established social bias taxonomies, VideoBiasEval employs an event-based prompting strategy to disentangle semantic content (actions and contexts) from actor attributes (gender and ethnicity). It further introduces multi-granular metrics to evaluate (1) overall ethnicity bias, (2) gender bias conditioned on ethnicity, (3) distributional shifts in social attributes across model variants, and (4) the temporal persistence of bias within videos. Using this framework, we conduct the first end-to-end analysis connecting biases in human preference datasets, their amplification in reward models, and their propagation through alignment-tuned video diffusion models. Our results reveal that alignment tuning not only strengthens representational biases but also makes them temporally stable, producing smoother yet more stereotyped portrayals. These findings highlight the need for bias-aware evaluation and mitigation throughout the alignment process to ensure fair and socially responsible video generation.

视频生成扩散模型社会偏见对齐训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。