首个评测视频隐喻理解能力的基准,推动模型从看画面到懂深意
ViMU: Benchmarking Video Metaphorical Understanding

- 构建首个系统性评估视频隐喻理解的基准数据集
- 要求模型基于多模态证据推理隐含意义,不依赖提示线索
- 适合关注视频深层语义、跨文化理解的研究者
任何新媒介一旦出现,其承载的信息通常包含两个层面:显性内容与隐含意义。视频技术普及后,不仅用于记录视觉信息,更成为传递情感、态度与社会意义的重要载体。许多视频的真正含义并不在画面本身,而藏于语境、表达风格及观众的社会经验中,体现为幽默、讽刺、批评等隐性内容,且不同文化背景解读差异显著。然而,当前多数视频理解模型仍聚焦于物体、动作或时间关系等字面理解,缺乏对隐喻、反讽与社会意义的系统分析能力。为此,我们提出ViMU——首个专门评估前沿模型视频隐含理解能力的基准。ViMU通过无提示设计的开放式与选择题,检验模型能否基于多模态证据推断隐含意义,确保所有关键线索均需模型自主发现。
原文摘要 · Abstract (English)
Any new medium, once it emerges, is used for more than the transmission of overt content alone. The information it carries typically operates on two levels: one is the content directly presented, while the other is the subtext beneath it-the implicit ideas and intentions the creator seeks to convey through the medium. Likewise, since video technologies became widely adopted, video has served not only as a powerful tool for recording and communicating visual information, but also as a vehicle for emotions, attitudes, and social meanings that are often difficult to articulate explicitly. Thus, the true meaning of many videos does not reside solely in what is shown on screen; it is often embedded in context, style of expression, and the viewer's social experience. Some forms of such video subtext are humorous, while others carry irony, mockery, or criticism. These implicit meanings can also be interpreted very differently across cultural backgrounds and social groups. However, most existing video understanding models still focus primarily on literal visual comprehension, such as recognizing objects, actions, or temporal relations, and lack a systematic ability to understand the metaphorical, ironic, and social meanings embedded in videos. To bridge this gap, we introduce ViMU, the first benchmark designed to systematically evaluate the subtext understanding capabilities of frontier models in videos. ViMU assesses whether video understanding models can go beyond literal perception to infer implicit meaning while grounding their interpretations in multimodal evidence and answering both open-ended and multiple-choice questions. Importantly, all questions are designed to be hint-free, ensuring that no key evidence is disclosed to models before answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。