研究生成模型间通信如何放大性别和年龄偏见
Investigating Associational Biases in Inter-Model Communication of Large Generative Models
- 通过图像生成与描述交替的通信链,追踪偏见传播路径
- 发现情感表达中女性化倾向增强,动作描述偏向年轻化
- 揭示模型依赖背景等无关视觉特征进行判断
生成式AI中的社会偏见不仅表现为性能差异,还体现为关联性偏见——即模型在无明确人口统计信息的情况下,仍会学习并重现概念与人群之间的刻板关联(如将医生与男性关联)。这种偏见可能在模型间通信链中持续传播并加剧,尤其在以人为中心的感知任务(如行为识别与情绪预测)中,导致错误推断与不平等对待。本研究聚焦人类行为与情感表达,采用RAF-DB与PHASE数据集,分析图像生成与描述交替过程中产生的代际与性别分布漂移。结果表明,动作与情绪表示均出现向更年轻、更多女性化方向的系统性漂移;部分预测依赖于背景或发型等无关视觉区域,而非身体或面部等关键线索。此外,验证了这些偏见会影响下游行为预测。最后提出从数据、训练到部署的多层级缓解策略,强调在人机交互系统中需建立严格防护机制。
原文摘要 · Abstract (English)
Social bias in generative AI can manifest not only as performance disparities but also as associational bias, whereby models learn and reproduce stereotypical associations between concepts and demographic groups, even in the absence of explicit demographic information (e.g., associating doctors with men). These associations can persist, propagate, and potentially amplify across repeated exchanges in inter-model communication pipelines, where one generative model's output becomes another's input. This is especially salient for human-centred perception tasks, such as human activity recognition and affect prediction, where inferences about behaviour and internal states can lead to errors or stereotypical associations that propagate into unequal treatment. In this work, focusing on human activity and affective expression, we study how such associations evolve within an inter-model communication pipeline that alternates between image generation and image description. Using the RAF-DB and PHASE datasets, we quantify demographic distribution drift induced by model-to-model information exchange and assess whether these drifts are systematic using an explainability pipeline. Our results reveal demographic drifts toward younger representations for both actions and emotions, as well as toward more female-presenting representations, primarily for emotions. We further find evidence that some predictions are supported by spurious visual regions (e.g., background or hair) rather than concept-relevant cues (e.g., body or face). We also examine whether these demographic drifts translate into measurable differences in downstream behaviour, i.e., while predicting activity and emotion labels. Finally, we outline mitigation strategies spanning data-centric, training and deployment interventions, and emphasise the need for careful safeguards when deploying interconnected models in human-centred AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。