用多个专家模型分解评估大模型回答,提升可信度。
Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs
- 将复杂回答拆解为原子任务,分派给擅长的专家模型评估。
- 构建新偏好数据集DGPref,使模型在信任度上显著提升。
- 基于开源模型,适合研究可信AI与多模态对齐的学者。
多模态大语言模型(MLLMs)可通过对齐人类偏好来增强可信度。由于人工标注偏好成本高,现有方法使用评估模型自动构建偏好数据集,但面对长且复合的响应,单一评估模型难以覆盖所有推理需求。此外,多数方法依赖闭源评估模型。为此,我们提出DecompGen框架,采用一组开源专家模型的集成方案。DecompGen将每个响应分解为原子验证任务,并将每项任务分配给最合适的专家模型生成细粒度评估。利用DecompGen反馈自动构建偏好数据集DGPref。通过偏好学习对齐DGPref的MLLMs在可信度上获得提升,证明了DecompGen的有效性。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) can enhance trustworthiness by aligning with human preferences. As human preference labeling is laborious, recent works employ evaluation models for assessing MLLMs' responses, using the model-based assessments to automate preference dataset construction. This approach, however, faces challenges with MLLMs' lengthy and compositional responses, which often require diverse reasoning skills that a single evaluation model may not fully possess. Additionally, most existing methods rely on closed-source models as evaluators. To address limitations, we propose DecompGen, a decomposable framework that uses an ensemble of open-sourced expert models. DecompGen breaks down each response into atomic verification tasks, assigning each task to an appropriate expert model to generate fine-grained assessments. The DecompGen feedback is used to automatically construct our preference dataset, DGPref. MLLMs aligned with DGPref via preference learning show improvements in trustworthiness, demonstrating the effectiveness of DecompGen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。