首个针对隐喻视频理解的基准与增强方法,提升大模型认知能力。
MetaphorVU: Towards Metaphorical Video Understanding

- 构建首个隐喻视频理解基准MetaphorVU-Bench,系统评估模型表现。
- 现有多模态大模型在隐喻理解上远低于人类水平,主因跨域映射缺陷。
- 提出MetaphorBoost框架,利用知识图谱增强推理,显著提升性能。
隐喻视频广泛存在于各类现实场景中,用于传达复杂概念,其理解通常需要高阶认知能力。目前对隐喻视频理解缺乏系统研究,不仅限制了多模态大模型(MLLMs)的实际应用,也阻碍了对其高阶认知能力的全面评估。为此,我们提出了首个系统且全面的隐喻视频理解基准MetaphorVU-Bench。实验发现,当前MLLMs在准确理解隐喻视频方面表现不佳,远落后于人类水平,主要归因于跨域映射能力缺陷。基于此,我们构建了隐喻知识图谱作为映射增强,并提出MetaphorBoost——一种推理时增强框架,在多个任务上实现持续性能提升。本研究的基准、分析与方法为推进MLLMs在高阶认知能力方面的发展提供了重要参考。
原文摘要 · Abstract (English)
Metaphorical videos are prevalent across various real-world scenarios to convey complex ideas, and understanding them typically requires high-order cognitive capabilities. The lack of systematic studies on metaphorical video understanding not only constrains the real-world applicability of MLLMs but also impedes the thorough assessment of their high-order cognitive capabilities. To bridge this gap, we propose MetaphorVU-Bench, the first systematic and comprehensive benchmark dedicated to metaphorical video understanding. Through experiments, we find current MLLMs struggle with accurate metaphorical video understanding, lagging far behind human level, primarily due to defective cross-domain mapping. Motivated by this finding, we construct a metaphor knowledge graph as mapping augmentation and propose MetaphorBoost, an inference-time enhancement framework achieving consistent performance improvement. Our benchmark, analysis, and method provide useful insights and a foundation for future research on advancing MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。