arXiv:2509.08777cs.CVcs.CL2025-09ICCV被引 3

用视觉特征动态调整提示权重,让图文生成评估更准确可靠

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles

  • 基于图像聚类的贝叶斯提示集成,根据图像特征动态分配提示权重
  • 在HPSv2和MJBench上提升判断准确率,显著改善评估结果的校准度
  • 适合需要高可信度图文生成评估的研究者与工业应用

多模态大语言模型(MLLM)正被广泛用于评估文本到图像(TTI)生成系统,基于视觉与文本上下文提供自动化判断。然而,这些“评判模型”常存在偏见、过度自信以及跨图像领域表现不一致的问题。尽管提示集成在纯文本场景中表现出色,但我们的实验表明,标准集成方法在TTI任务中难以泛化。为此,我们提出一种新的多模态感知方法——多模态贝叶斯提示集成(MMB)。该方法通过图像聚类增强贝叶斯提示集成,使评判模型能根据样本的视觉特征动态分配提示权重。实验显示,MMB在成对偏好判断中提升了准确性,并显著增强了校准性,更真实反映模型不确定性。在两个TTI基准测试集HPSv2和MJBench上的评估结果表明,MMB在与人类标注对齐及跨多样图像内容的校准性方面均优于现有基线。研究强调了为评判模型设计多模态特异性策略的重要性,为大规模可靠TTI评估指明了新方向。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are increasingly used to evaluate text-to-image (TTI) generation systems, providing automated judgments based on visual and textual context. However, these "judge" models often suffer from biases, overconfidence, and inconsistent performance across diverse image domains. While prompt ensembling has shown promise for mitigating these issues in unimodal, text-only settings, our experiments reveal that standard ensembling methods fail to generalize effectively for TTI tasks. To address these limitations, we propose a new multimodal-aware method called Multimodal Mixture-of-Bayesian Prompt Ensembles (MMB). Our method uses a Bayesian prompt ensemble approach augmented by image clustering, allowing the judge to dynamically assign prompt weights based on the visual characteristics of each sample. We show that MMB improves accuracy in pairwise preference judgments and greatly enhances calibration, making it easier to gauge the judge's true uncertainty. In evaluations on two TTI benchmarks, HPSv2 and MJBench, MMB outperforms existing baselines in alignment with human annotations and calibration across varied image content. Our findings highlight the importance of multimodal-specific strategies for judge calibration and suggest a promising path forward for reliable large-scale TTI evaluation.

图文生成模型校准提示工程多模态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。