arXiv:2607.15740cs.CVcs.AI2026-07

用轻量模型提升文生图评估的文化准确性,速度快10倍。

Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling

论文配图:Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
图 1 · 摘自论文原文
  • 基于42亿参数多模态大模型,引入隐式文化探测与跳连注意力机制
  • 在CulturalFrames数据集上达到82.12%配对准确率,相关系数超同类方法
  • 无需自回归生成,单次评估仅需0.21秒,适合实时偏好优化

随着文生图系统快速发展,评估生成内容的文化真实性对公平可信的生成AI至关重要。现有评估指标和多模态评判器常依赖视觉-语义表征,难以捕捉隐含文化规范,导致偏好判断偏差并忽略细微文化线索。此外,基于VQA的评估器通常依赖自回归文本生成,限制了其在实时奖励建模中的可扩展性。为此,我们提出一种基于轻量级42亿参数多模态大模型(MLLM)的隐式文化对齐奖励模型。框架融合隐式文化探测与跳连交叉注意力(SkipCA)机制,使晚期语义特征可直接关注早期视觉表示,更好地保留文化显著细节。在CulturalFrames基准上3,323组精心筛选的图像对评估中,本方法达到82.12%的配对准确率,皮尔逊与肯德尔相关系数分别为0.585和0.412,优于代表性视觉语言指标和基于MLLM的评估器。同时,通过跳过自回归文本生成,单次评估耗时0.21秒,较标准VQA评估器提速10倍。结果表明,该奖励模型可为强化学习人类反馈(RLHF)与直接偏好优化(DPO)等偏好优化流程提供高效且文化敏感的标量信号。

原文摘要 · Abstract (English)

As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI. Existing T2I evaluation metrics and multimodal judges often rely on visual-semantic representations that underrepresent implicit cultural norms, leading to biased preference judgments and the omission of fine-grained cultural cues. In addition, visual question answering (VQA)-based evaluators typically depend on autoregressive text generation, which limits their scalability for real-time reward modeling. To address these limitations, we introduce an Implicit Cultural Alignment Reward Model built upon a lightweight 4.2-billion-parameter Multimodal Large Language Model (MLLM). Our framework integrates an Implicit Cultural Probe with a Skip-connection Cross-Attention (SkipCA) mechanism, enabling late-stage semantic features to directly attend to early-stage visual representations and better preserve culturally salient details. Evaluations on 3,323 challenging and carefully curated image pairs from the CulturalFrames benchmark show that our approach achieves 82.12% pairwise accuracy, with Pearson and Kendall correlation coefficients of 0.585 and 0.412, respectively, outperforming representative vision-language metrics and MLLM-based evaluators. Moreover, by bypassing autoregressive text generation, our model processes each evaluation in 0.21 seconds under our local inference setup, achieving a $10\times$ speedup over standard VQA-based evaluators. These results suggest that the proposed reward model can provide an efficient and culturally aware scalar signal for preference optimization pipelines such as Reinforcement Learning from Human Feedback and Direct Preference Optimization. Additional resources are available on our project page at https://bensonch1214.github.io/Implicit_Cultural_Alignment/.

文生图文化对齐奖励模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。