arXiv:2505.23065cs.CL2025-05被引 3

构建社交平台多模态评测基准,评估模型理解图文内容能力

SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services

  • 设计8类真实社交场景任务,覆盖图文理解与推荐等
  • 含4001组精心标注的图文问答对,支持多种题型
  • 适合研究多模态大模型在社交应用中的表现与优化

随着社交网络服务(SNS)中视觉与文本内容融合日益加深,评估大型语言模型(LLMs)的多模态能力对提升用户体验、内容理解和平台智能至关重要。现有评测基准多聚焦文本任务,缺乏对现代SNS中普遍存在的多模态场景覆盖。本文提出SNS-Bench-VL,一个全面的多模态基准,用于评估视觉-语言大模型在真实社交媒体场景下的表现。该基准涵盖8类多模态任务,包括笔记理解、用户参与度分析、信息检索和个性化推荐,包含4001个精心设计的图文问答对,涵盖单选、多选和开放问答题型。我们评估了25个以上前沿多模态LLM,分析其在各任务上的表现。结果揭示了模型在多模态社交语境理解方面仍存在显著挑战。我们希望SNS-Bench-VL能推动下一代面向社交服务的鲁棒、上下文感知且符合人类对齐的多模态智能研究。

原文摘要 · Abstract (English)

With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial for enhancing user experience, content understanding, and platform intelligence. Existing benchmarks primarily focus on text-centric tasks, lacking coverage of the multimodal contexts prevalent in modern SNS ecosystems. In this paper, we introduce SNS-Bench-VL, a comprehensive multimodal benchmark designed to assess the performance of Vision-Language LLMs in real-world social media scenarios. SNS-Bench-VL incorporates images and text across 8 multimodal tasks, including note comprehension, user engagement analysis, information retrieval, and personalized recommendation. It comprises 4,001 carefully curated multimodal question-answer pairs, covering single-choice, multiple-choice, and open-ended tasks. We evaluate over 25 state-of-the-art multimodal LLMs, analyzing their performance across tasks. Our findings highlight persistent challenges in multimodal social context comprehension. We hope SNS-Bench-VL will inspire future research towards robust, context-aware, and human-aligned multimodal intelligence for next-generation social networking services.

多模态评测基准社交网络大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。