arXiv:2412.08687cs.CV2024-12CVPR被引 27

构建23万条真实用户与视觉语言模型对话数据集,用于评估模型真实交互表现。

VisionArena: 230K Real World User-VLM Conversations with Preference Labels

  • 从真实用户平台收集23万次对话,涵盖73000名用户与45种模型
  • 基于用户偏好投票,发现模型在空间推理和规划任务上表现差
  • 用该数据集微调模型,性能超越Llava-Instruct-158K,MMMU提升17点

随着视觉语言模型(VLMs)的广泛应用与能力提升,亟需能捕捉真实用户-模型交互的基准。为此,我们构建了VisionArena,一个包含23万条真实世界用户与VLM对话的数据集。数据来自Chatbot Arena——一个开源平台,用户在此与VLM互动并提交偏好投票。VisionArena涵盖7.3万唯一用户、45个VLM和138种语言。数据集包含三个子集:VisionArena-Chat(20万条单轮或多轮对话)、VisionArena-Battle(3万条匿名双模型对比对话及用户偏好投票)、VisionArena-Bench(500个多样化用户提示的自动基准,可高效逼近实时模型排名)。我们分析了用户提问类型、响应风格对偏好影响,以及模型常见失败场景。发现开放任务如图像描述和幽默理解高度依赖风格,当前VLM在空间推理与规划任务上表现不佳。最后,将同一基础模型在VisionArena-Chat上微调后,在MMMU上获得17点提升,WildVision上提升46点,优于Llava-Instruct-158K。

原文摘要 · Abstract (English)

With the growing adoption and capabilities of vision-language models (VLMs) comes the need for benchmarks that capture authentic user-VLM interactions. In response, we create VisionArena, a dataset of 230K real-world conversations between users and VLMs. Collected from Chatbot Arena - an open-source platform where users interact with VLMs and submit preference votes - VisionArena spans 73K unique users, 45 VLMs, and 138 languages. Our dataset contains three subsets: VisionArena-Chat, 200k single and multi-turn conversations between a user and a VLM; VisionArena-Battle, 30K conversations comparing two anonymous VLMs with user preference votes; and VisionArena-Bench, an automatic benchmark of 500 diverse user prompts that efficiently approximate the live Chatbot Arena model rankings. Additionally, we highlight the types of question asked by users, the influence of response style on preference, and areas where models often fail. We find open-ended tasks like captioning and humor are highly style-dependent, and current VLMs struggle with spatial reasoning and planning tasks. Lastly, we show finetuning the same base model on VisionArena-Chat outperforms Llava-Instruct-158K, with a 17-point gain on MMMU and a 46-point gain on the WildVision benchmark. Dataset at https://huggingface.co/lmarena-ai

视觉语言模型对话数据集用户偏好模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。