arXiv:2509.11620cs.CLcs.CY2025-09EMNLP被引 11

评测多模态大模型在个性化审美评估中的偏见与对齐度

AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment

  • 构建基准AesBiasBench,从刻板印象偏见和人类偏好对齐两维度评估
  • 小模型偏见更明显,大模型更贴近真实人类审美,身份信息加剧情感判断偏差
  • 适用于关注视觉-语言任务中公平性与可解释性的研究者与开发者

多模态大语言模型(MLLMs)正被广泛用于个性化图像审美评估(PIAA),作为专家评价的可扩展替代方案。然而,其预测可能受性别、年龄、教育等人口统计因素影响而产生微妙偏见。本文提出AesBiasBench,一个用于评估MLLMs在两个互补维度上的基准:(1) 刻板印象偏见,通过测量不同人口群体间审美评价的差异来量化;(2) 模型输出与真实人类审美偏好之间的对齐程度。该基准涵盖三个子任务(审美感知、评估、共情),引入结构化指标(IFD、NRD、AAS)以同时评估偏见与对齐。我们评估了19个MLLMs,包括专有模型(如GPT-4o、Claude-3.5-Sonnet)和开源模型(如InternVL-2.5、Qwen2.5-VL)。结果表明,较小模型表现出更强的刻板印象偏见,而较大模型更贴近人类偏好。在情感判断中加入身份信息往往加剧偏见。这些发现强调了在主观视觉-语言任务中采用身份敏感评估框架的重要性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are increasingly applied in Personalized Image Aesthetic Assessment (PIAA) as a scalable alternative to expert evaluations. However, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education. In this work, we propose AesBiasBench, a benchmark designed to evaluate MLLMs along two complementary dimensions: (1) stereotype bias, quantified by measuring variations in aesthetic evaluations across demographic groups; and (2) alignment between model outputs and genuine human aesthetic preferences. Our benchmark covers three subtasks (Aesthetic Perception, Assessment, Empathy) and introduces structured metrics (IFD, NRD, AAS) to assess both bias and alignment. We evaluate 19 MLLMs, including proprietary models (e.g., GPT-4o, Claude-3.5-Sonnet) and open-source models (e.g., InternVL-2.5, Qwen2.5-VL). Results indicate that smaller models exhibit stronger stereotype biases, whereas larger models align more closely with human preferences. Incorporating identity information often exacerbates bias, particularly in emotional judgments. These findings underscore the importance of identity-aware evaluation frameworks in subjective vision-language tasks.

多模态模型偏见评估审美生成公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。