用少量文本推理数据训练,实现跨模态零样本评估。
Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
- 基于文本推理构建通用决策模式,实现跨模态迁移。
- 仅用少量文本数据,性能优于主流商业API和专用模型。
- 适合资源稀缺领域如分子生成的高效评估。
人类生成的奖励信号对对齐生成模型与人类偏好至关重要,既用于训练也用于推理阶段评估。尽管将大语言模型(LLMs)作为代理评估者(即 LLM-as-a-Judge)可显著降低人工标注成本,但它们通常需要大量特定模态的训练数据,且在多样化多模态任务中泛化能力差。本文提出 Flex-Judge,一种基于推理引导的多模态评估模型,仅需极少文本推理数据即可在多种模态和评估格式间稳健泛化。核心思路是:结构化的文本推理解释天然包含可迁移的决策模式,能有效推广至图像、视频等多模态判断。实验证明,尽管训练数据远少于现有方法,Flex-Judge 在多项任务上表现媲美甚至超越先进商用 API 和深度训练的多模态评估器。尤其在分子等缺乏完整评估基准的模态中展现出广泛影响,凸显其在资源受限场景下的实用价值。本框架表明,以推理为基础的文本监督是一种高效、低成本的替代方案,大幅推进了可扩展的多模态模型作为评估者的发展。
原文摘要 · Abstract (English)
Human-generated reward signals are critical for aligning generative models with human preferences, guiding both training and inference-time evaluations. While large language models (LLMs) employed as proxy evaluators, i.e., LLM-as-a-Judge, significantly reduce the costs associated with manual annotations, they typically require extensive modality-specific training data and fail to generalize well across diverse multimodal tasks. In this paper, we propose Flex-Judge, a reasoning-guided multimodal judge model that leverages minimal textual reasoning data to robustly generalize across multiple modalities and evaluation formats. Our core intuition is that structured textual reasoning explanations inherently encode generalizable decision-making patterns, enabling an effective transfer to multimodal judgments, e.g., with images or videos. Empirical results demonstrate that Flex-Judge, despite being trained on significantly fewer text data, achieves competitive or superior performance compared to state-of-the-art commercial APIs and extensively trained multimodal evaluators. Notably, Flex-Judge presents broad impact in modalities like molecule, where comprehensive evaluation benchmarks are scarce, underscoring its practical value in resource-constrained domains. Our framework highlights reasoning-based text supervision as a powerful, cost-effective alternative to traditional annotation-intensive approaches, substantially advancing scalable multimodal model-as-a-judge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。