让多模态模型更靠谱地评判视觉内容,避免被文字误导。
Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling

- 用视觉扰动生成反事实数据,让模型学会依赖真实感知。
- 在多个评测中显著提升判断准确性和一致性,接近人类水平。
- 适合需要可靠自动评估的科研与工业场景,如图像理解评测。
近期多模态大模型虽具强大推理能力,但作为自动化评估者时存在关键缺陷:当视觉信息与文本线索冲突时,其倾向奖励看似合理而非感知正确的答案。我们识别并系统分析了这一现象,称为「感知判断偏差」。通过受控的视觉扰动实验发现,现有多模态判官常锚定于文本回应,而非自身视觉感知,导致评价不一致且不可验证。为此,我们构建了感知扰动评判数据集(Perceptually Perturbed Judgment Dataset),通过最小编辑生成反事实响应,分离感知错误并实现可验证监督。基于该数据集,提出统一训练框架,结合结构化GRPO奖励与批量排序目标,无需显式成对标签即可实现连贯全局排序。在多个MLLM-as-a-Judge基准上实验表明,该方法显著提升感知保真度、排序一致性及与人类评估的一致性。结果建立了一条可扩展、通用的多模态判官训练路径,使其感知根植、可解释且抗视觉-推理冲突。
原文摘要 · Abstract (English)
Recent multimodal large language models have demonstrated strong reasoning ability, yet their reliability as automated evaluators remains limited by a critical weakness: when visual evidence conflicts with textual cues, MLLM judges tend to reward plausible narratives over perceptually correct answers. We identify and systematically analyze this phenomenon, which we term Perceptual Judgment Bias. Through controlled visual perturbations, existing multimodal judges frequently anchor on the response text instead of their own visual perception, leading to inconsistent and non-verifiable evaluations. To address this issue, we introduce the Perceptually Perturbed Judgment Dataset, which constructs minimally edited counterfactual responses that isolate perceptual errors and enable verifiable supervision. Building on this dataset, we develop a unified training framework that combines a structured GRPO-based reward with a batch-ranking objective, achieving coherent global ordering without explicit pairwise labels. Experiments across diverse MLLM-as-a-Judge benchmarks show that our approach substantially improves perceptual fidelity, ranking coherence, and alignment with human evaluation. Our results establish a scalable and generalizable pathway for training multimodal judges that are perceptually grounded, interpretable, and robust to visual-reasoning conflicts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。