评测模型理解社交媒体视频隐含意义的能力
Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

- 构建1000个带隐含叙事注释的短视频数据集,聚焦多模态语境推理
- 现有模型在解释视频真实意图上表现不佳,距人类理解有明显差距
- 适合研究多模态语义理解、讽刺识别与对话系统的人参考
社交媒体视频常传递超出画面、字幕或语音的深层含义。一个看似平淡的片段,可能因多模态线索与文化背景的互动而具有幽默、讽刺或反讽意味,这对视频-语言模型构成严峻挑战。本文提出DrivelHub+,一个包含1000个社交平台视频的基准数据集,每个视频配有由人工撰写的隐含叙事解释。不同于传统以识别或描述为主的视频理解任务,本研究聚焦上下文驱动的多模态推理。我们从两个角度评估当前视频-语言模型:解释任务中要求模型用自然语言阐明视频的语用理解;表示任务中采用推理即检索方法,测试模型表征是否能将视频与其对应隐含叙事准确匹配(视频到文本、文本到视频)。该基准为衡量多模态感知与语用理解之间的鸿沟提供了诊断性场景,追问当前模型能否超越‘看什么’,实现‘懂什么’。
原文摘要 · Abstract (English)
Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context, making such content a difficult test case for video-language models. In this paper, we introduce \textit{DrivelHub+}, a benchmark for evaluating whether models can infer the implicit, non-linear, and rhetorically layered meanings of social media videos that appear nonsensical on the surface but convey deliberate pragmatic meanings. DrivelHub+ consists of 1,000 videos collected from social media, each annotated with a human-written implicit narrative explanation. Unlike conventional video understanding tasks focused on recognition or description, we present a benchmark that targets contextual multimodal reasoning. We evaluate current video-language models from two perspectives: explanation, where models must explain the pragmatic comprehension of a video in natural language; and representation, where we adapt reasoning-as-retrieval to test whether model representations align videos with their corresponding implicit narratives in both video-to-text and text-to-video retrieval. Our benchmark provides a diagnostic setting for measuring the gap between multimodal perception and pragmatic comprehension, asking whether current models can move beyond describing what is shown to inferring what is meant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。