arXiv:2605.08175cs.CVcs.AI2026-05被引 1

构建音乐视频因果问答基准,测试视觉如何影响音乐结构。

KARMA-MV: A Benchmark for Causal Question Answering on Music Videos

论文配图:KARMA-MV: A Benchmark for Causal Question Answering on Music Videos
图 1 · 摘自论文原文
  • 用大模型自动生成3.7万道多选题,覆盖时序音视频线索。
  • 引入因果知识图谱,使模型在小规模下也能提升推理准确率。
  • 适合研究音视频因果推理、跨模态理解的学者和开发者。

尽管视频问答和跨模态理解取得进展,但音乐视频中视觉动态如何驱动音乐结构的因果推理仍被忽视。我们提出KARMA-MV,一个基于2,682个YouTube音乐视频的大规模多选题数据集,用于评估模型整合时序音视频线索并回答关于视觉到音乐影响的推理、预测与反事实问题的能力。不同于需人工标注的传统数据集,KARMA-MV利用大语言模型(LLM)进行可扩展生成与验证,共生成37,737道多选题。我们提出一种因果知识图谱(CKG)方法,通过结构化检索跨模态依赖关系,增强视觉-语言模型(VLM)表现。在先进VLM与LLM上的实验表明,引入CKG能持续提升性能,尤其对小型模型效果显著,证明显式因果结构对音乐视频理解的价值。KARMA-MV为超越相关性的因果音频-视觉理解提供了新基准。

原文摘要 · Abstract (English)

While significant progress has been made in Video Question Answering and cross-modal understanding, causal reasoning about how visual dynamics drive musical structure in music videos remains under-explored. We introduce KARMA-MV, a large-scale multiple-choice QA dataset derived from 2,682 YouTube music videos, designed to test models' ability to integrate temporal audio-visual cues and reason about visual-to-musical influence across reasoning, prediction, and counterfactual questions. Unlike traditional datasets requiring manual annotation, KARMA-MV leverages LLM reasoning for scalable generation and validation, yielding 37,737 MCQs. We propose a causal knowledge graph (CKG) approach that augments vision-language models (VLMs) with structured retrieval of cross-modal dependencies. Experiments on state-of-the-art VLMs and LLMs show consistent gains from CKG grounding -- especially for smaller models -- establishing the value of explicit causal structure for music-video reasoning. KARMA-MV provides a new benchmark for advancing causal audio-visual understanding beyond correlation.

因果推理音视频理解多模态问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。