发现统一多模态大模型生成图像存在显著性别种族偏见,提出定位修复策略。
On Fairness of Unified Multimodal Large Language Model for Image Generation
- 通过定位分析发现偏见主要来自语言模型组件。
- 实验证明该模型在理解任务中偏见小,但生成任务中偏见严重。
- 用合成数据训练平衡偏好模型,有效降低偏见且保持语义准确。
统一多模态大语言模型(U-MLLMs)在端到端视觉理解与生成中表现优异,但相比仅生成类模型(如Stable Diffusion),其统一能力可能引发新的偏见问题。本文对最新U-MLLMs进行基准测试,发现多数模型存在显著的性别与种族偏见。为深入理解并缓解此问题,我们提出“定位-修复”策略,审计各组件偏见来源。分析表明,偏见主要源自语言模型。更有趣的是,观察到“部分对齐”现象:理解任务中偏见微弱,但生成任务中偏见仍显著。为此,我们提出一种新型平衡偏好模型,结合合成数据以均衡人口分布。实验表明,该方法在降低偏见的同时保持了良好的语义保真度。研究呼吁未来需采用更全面的视角审视与去偏策略。
原文摘要 · Abstract (English)
Unified multimodal large language models (U-MLLMs) have demonstrated impressive performance in visual understanding and generation in an end-to-end pipeline. Compared with generation-only models (e.g., Stable Diffusion), U-MLLMs may raise new questions about bias in their outputs, which can be affected by their unified capabilities. This gap is particularly concerning given the under-explored risk of propagating harmful stereotypes. In this paper, we benchmark the latest U-MLLMs and find that most exhibit significant demographic biases, such as gender and race bias. To better understand and mitigate this issue, we propose a locate-then-fix strategy, where we audit and show how the individual model component is affected by bias. Our analysis shows that bias originates primarily from the language model. More interestingly, we observe a "partial alignment" phenomenon in U-MLLMs, where understanding bias appears minimal, but generation bias remains substantial. Thus, we propose a novel balanced preference model to balance the demographic distribution with synthetic data. Experiments demonstrate that our approach reduces demographic bias while preserving semantic fidelity. We hope our findings underscore the need for more holistic interpretation and debiasing strategies of U-MLLMs in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。