arXiv:2510.22827cs.CVcs.LG2025-10

用带拒答机制的多模态模型评估文生图公平性,更贴近人类判断。

FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models

  • 用指令跟随的多模态大模型作为结构化裁判,支持拒答和证据说明。
  • 在4个属性预测与3个对齐评测中优于或补充现有方法。
  • 适合关注图像生成公平性与可解释性的研究者与开发者。

评估文本到图像(T2I)系统不仅需判断图像是否匹配提示,还需考察社会敏感属性是否忠实呈现且无过度推断。现有自动化评估工具多依赖人脸识别器或图文对比相似度,反馈有限且常强制输出,即使视觉证据模糊或缺失。对于宗教、残疾等属性,其线索可能具情境性、间接性或故意不显,此类工具易遗漏人类审阅者能发现的缺陷。本文提出 extsc{FairJudge},一种支持拒答的多模态评估协议,使用指令遵循的多模态大模型进行社会属性预测、职业定位及提示-图像对齐判断。该协议限制输出为封闭标签集,要求可见证据理由,允许在线索不足时选择“未指定”,并将评分映射至 $[-1,1]$。这些约束使大模型评估从开放式判断变为可解析、可审计的流程。在四个属性预测基准与三个职业/对齐基准上, extsc{FairJudge} 表现优于或补充了 CLIP、DeepFace、VIEScore 与 VQAScore。消融实验表明,封闭标签、拒答机制与证据报告是可靠性核心。我们还引入 extsc{DIVERSIFY} 与 extsc{DIVERSIFY-Professions},两个富含上下文的资源,用于评估超越可见面部或典型符号的社会表征与职业定位。代码、提示、数据集、解析日志与逐图裁判结果均已公开,支持可复现审计。

原文摘要 · Abstract (English)

Evaluating text-to-image (T2I) systems requires judging not only whether an image matches a prompt, but also whether socially salient attributes are represented faithfully and without unsupported inference. Existing automated evaluators typically rely on face-centric recognizers or contrastive image--text similarity, which provide limited diagnostic feedback and often force predictions even when visual evidence is ambiguous or absent. For fairness-sensitive attributes such as religion and disability, where cues may be contextual, indirect, or intentionally unspecified, these evaluators can therefore miss failure modes that careful human reviewers would notice. We introduce \textsc{FairJudge}, an abstention-aware evaluation protocol that uses instruction-following multimodal LLMs as structured judges for social-attribute prediction, profession grounding, and prompt--image alignment. The protocol constrains outputs to closed label sets, requires visible-evidence rationales, supports an explicit \textsc{unspecified} decision when cues are insufficient, and maps rubric-based alignment judgments to $[-1,1]$. These constraints turn MLLM judging from open-ended assessment into a parseable, auditable evaluation procedure. Across four attribute-prediction benchmarks and three profession/alignment benchmarks, \textsc{FairJudge} outperforms or complements CLIP, DeepFace, VIEScore, and VQAScore. Ablations show that closed labels, abstention, and evidence reporting are central to reliability. We further introduce \textsc{DIVERSIFY} and \textsc{DIVERSIFY-Professions}, two context-rich resources for evaluating social representation and profession grounding beyond face-visible or iconic cues. We release code, prompts, datasets, parser logs, and per-image judge outputs to support reproducible auditing.

公平性评估多模态大模型文生图可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。