多模态大模型在跨文化评判中存在校准与方向偏差,影响公平性。
Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity

- 构建跨中美文化对比的626项图文评估集,发现人类评判存在显著文化差异。
- 六种大模型普遍存在正向底限校准失效和默认单一文化倾向的问题。
- 建议分池报告对齐结果,将跨池分歧视为评判模型的固有属性。
MLLM作为评判者通常依赖与人类标注的一致性验证,但在文化异质的人类群体中该指标失效。本文提出VOIR DIRE,一个包含626个跨文化配对图文样本的多模态基准,覆盖美中两国在食物、时尚与建筑领域的语境,其组内标注者可靠性分别为a = 0.86/0.74,但跨组评价相关性极低(Q1 r = -0.12)。六种MLLM的偏差可分解为两类:正向底限校准失败(尺度压缩)与方向性偏差(默认单一文化规范)。在争议样本被用于分割两组的测试中,模型机械地支持更宽松的中国解读;角色提示部分恢复校准,但方向残留仍存,表明倾斜非仅由尺度压缩导致。参考池上下文示范反而加深方向残留并抬高高分端,而非恢复低分端使用。模型来源带来约0.10 MAE的小幅加性偏倚,且在示范下基本不变。建议分别报告与各参考池的对齐情况,并将跨池分歧视作评判模型的属性。
原文摘要 · Abstract (English)
MLLM-as-a-Judge is conventionally validated by agreement with human annotations, but this metric is undefined when the human pool is culturally heterogeneous. We introduce VOIR DIRE, a multimodal benchmark of 626 culturally paired image--prompt artifacts spanning U.S. and mainland Chinese contexts across food, fashion, and architecture, with annotator pools that are within-pool reliable (a = 0.86/0.74) but cross-pool divergent on evaluation (Q1 r = -0.12). Across six MLLMs, the bias decomposes into two failures: a positivity-floor calibration failure (compressed scale use) and an orientation failure (default to one cultural norm). On this corpus, where contested items are sampled to split the two pools, the floor mechanically validates the more-permissive Chinese reading; persona prompting partially recovers calibration, but the orientation residual survives, evidence the tilt is not reducible to scale compression. Reference-pool in-context demonstrations deepen the orientation residual and inflate the high end rather than restoring use of the low end. Model origin adds a small additive tilt (~0.10 MAE) that is approximately invariant under demonstration. We recommend reporting alignment against each reference pool separately and treating cross-pool divergence as a judge property.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。