arXiv:2411.11479cs.CL2024-11ACL

用社交媒体视频评测视觉语言模型的价值偏好与角色扮演能力

Value-Spectrum: Quantifying Preferences of Vision-Language Models via Value Decomposition in Social Media Contexts

论文配图:Value-Spectrum: Quantifying Preferences of Vision-Language Models via Value Decomposition in Social Media Contexts
图 1 · 摘自论文原文
  • 构建包含5万+短视频的向量数据库,模拟真实浏览场景
  • 发现不同模型在价值观任务中表现差异显著,部分可主动切换人格角色
  • 适合研究模型价值观对齐与多角色生成的学者使用

视觉语言模型(VLM)的发展拓展了多模态应用边界,但现有评估多局限于功能任务,忽视人格特质与人类价值观等抽象维度。为此,我们提出Value-Spectrum,一个基于施瓦茨价值维度的视觉问答(VQA)基准,用于评估VLM在价值观层面的表现。我们设计了一个VLM代理流程,模拟视频浏览行为,构建了一个包含超过5万条短视频的向量数据库,数据源来自TikTok、YouTube Shorts和Instagram Reels,覆盖家庭、健康、兴趣、社会、科技等多个主题,时间跨度达数月。在Value-Spectrum上的测试揭示了不同VLM在处理价值导向内容时的显著差异。此外,我们还探索了在明确提示下VLM代理是否能采用特定人格,结果表明模型具备一定的角色扮演适应能力。这些发现表明Value-Spectrum可作为全面评估VLM价值偏好及角色模拟能力的基准。完整代码与数据已开源:https://github.com/Jeremyyny/Value-Spectrum。

原文摘要 · Abstract (English)

The recent progress in Vision-Language Models (VLMs) has broadened the scope of multimodal applications. However, evaluations often remain limited to functional tasks, neglecting abstract dimensions such as personality traits and human values. To address this gap, we introduce Value-Spectrum, a novel Visual Question Answering (VQA) benchmark aimed at assessing VLMs based on Schwartz's value dimensions that capture core human values guiding people's preferences and actions. We design a VLM agent pipeline to simulate video browsing and construct a vector database comprising over 50,000 short videos from TikTok, YouTube Shorts, and Instagram Reels. These videos span multiple months and cover diverse topics, including family, health, hobbies, society, technology, etc. Benchmarking on Value-Spectrum highlights notable variations in how VLMs handle value-oriented content. Beyond identifying VLMs' intrinsic preferences, we also explore the ability of VLM agents to adopt specific personas when explicitly prompted, revealing insights into the adaptability of the model in role-playing scenarios. These findings highlight the potential of Value-Spectrum as a comprehensive evaluation set for tracking VLM preferences in value-based tasks and abilities to simulate diverse personas. The complete code and data are available at: https://github.com/Jeremyyny/Value-Spectrum.

视觉语言模型价值观评估角色扮演多模态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。