构建动态偏好数据集,评估视觉语言模型实时适应用户偏好的能力。
A Dataset for Dynamic Human Preferences for Vision Language Models

- 设计自动化流程生成动态多模态偏好数据
- 首次评估大模型在推理时理解实时偏好的能力
- 适合研究人机交互与个性化模型的学者
随着视觉语言模型(VLMs)在人机交互场景中的广泛应用,评估其适应不同用户实时偏好的能力至关重要。尽管近年来涌现出大量视觉语言基准测试,但它们主要关注静态能力与普遍性偏好,这些偏好通常来自大规模训练数据。本文提出一个新基准,用于评估VLMs理解动态人类偏好的能力——即在推理时以上下文方式传递的偏好。我们提供了自动化数据生成管道,涵盖图像依赖性变化的多种情形,并构建了一个动态多模态人类偏好数据集。此外,对当前主流模型在该基准上的评估结果也一并给出。
原文摘要 · Abstract (English)
Given the increased adoption of Vision Language Models (VLMs) in human-interactive settings, it is important that we evaluate how well these models can adapt to real-time preferences for different users. While an increasing number of vision-language benchmarks have recently been introduced, they focus largely on evaluating static capabilities and generally-held preferences learned from extensive training data. This work introduces a new benchmark for evaluating the ability of VLMs to understand dynamic human-preferences, i.e. preferences that are passed in-context at inference time. We provide an automated pipeline for generating this benchmark with variations on image dependence, a dynamic multi-modal human-preference dataset, and evaluations of state-of-the-art models on the novel benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。