构建首个视频幽默理解基准,测试多模态模型看图猜笑的能力
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
- 设计非语言短视频集合,仅靠视觉判断幽默
- 实验证明纯视觉理解幽默仍具挑战性
- 引入音效增强分析,适合研究多模态视频理解的学者
具备理解幽默能力的AI模型具有实际应用前景,例如提升人机交互的参与感。为评估和诊断多模态大语言模型(MLLMs)在幽默理解方面的能力,我们提出v-HUB,一个全新的视频幽默理解基准。v-HUB包含一组精心筛选的非语言短视频,反映真实场景中仅通过视觉线索即可感知幽默的情况。每个视频片段均配有丰富标注,支持多种评估任务与分析,包括首次研究环境音效对幽默感知的增强作用。为提升适用性,我们构建了开放式问答任务,使v-HUB可无缝集成至现有视频理解任务套件中。我们评估了涵盖专用Video-LLMs与原生处理音频的通用OmniLLMs的多种MLLMs,结果揭示了当前模型在仅依赖视觉线索理解幽默时仍面临显著困难。研究还表明,引入音频有助于提升视频幽默理解效果,凸显融合更丰富模态在复杂视频理解任务中的潜力。
原文摘要 · Abstract (English)
AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge and diagnose the capacity of multimodal large language models (MLLMs) for humor understanding, we introduce v-HUB, a novel video humor understanding benchmark. v-HUB comprises a curated collection of non-verbal short videos, reflecting real-world scenarios where humor can be appreciated purely through visual cues. We pair each video clip with rich annotations to support a variety of evaluation tasks and analyses, including a novel study of environmental sound that can enhance humor. To broaden its applicability, we construct an open-ended QA task, making v-HUB readily integrable into existing video understanding task suites. We evaluate a diverse set of MLLMs, from specialized Video-LLMs to versatile OmniLLMs that can natively process audio, covering both open-source and proprietary domains. The experimental results expose the difficulties MLLMs face in comprehending humor from visual cues alone. Our findings also demonstrate that incorporating audio helps with video humor understanding, highlighting the promise of integrating richer modalities for complex video understanding tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。