arXiv:2601.17645cs.SDcs.CL2026-01被引 3

测试大模型对网络声音视频的文化理解能力,发现其远不如人类。

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

  • 构建千余条音视频数据集,涵盖歌曲、语音、音效等多模态内容
  • 模型在无文字的音乐和音效上表现差,难以理解上下文与文化背景
  • 适合研究多模态认知、文化理解与模型评估的研究者参考

互联网音视频通过随时间变化的声音与动作传递意义,超越了文本所能表达的范围。为检验AI模型是否能理解此类信号中的人类文化语境,我们提出了AVMeme Exam,一个由人工标注的基准测试集,包含超过一千个具有代表性的网络声音与视频片段,涵盖语音、歌曲、音乐及音效。每个迷因(meme)均配有独特的问答对,评估从表面内容到上下文、情感、使用方式及世界知识等多个层面的理解。同时提供原始年份、转录文本、摘要和敏感性元数据。我们系统评估了当前最先进的多模态大模型(MLLMs),并与人类参与者对比。结果表明:现有模型在无文字的音乐和音效上表现不佳,难以进行情境化和文化化思考,相比表面内容理解存在明显差距。这一发现揭示了人类对齐的多模态智能中的关键缺口,呼吁开发能感知上下文与文化的模型。项目页面:avmemeexam.github.io/public

原文摘要 · Abstract (English)

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we introduce AVMeme Exam, a human-curated benchmark of over one thousand iconic Internet sounds and videos spanning speech, songs, music, and sound effects. Each meme is paired with a unique Q&A assessing levels of understanding from surface content to context and emotion to usage and world knowledge, along with metadata such as original year, transcript, summary, and sensitivity. We systematically evaluate state-of-the-art multimodal large language models (MLLMs) alongside human participants using this benchmark. Our results reveal a consistent limitation: current models perform poorly on textless music and sound effects, and struggle to think in context and in culture compared to surface content. These findings highlight a key gap in human-aligned multimodal intelligence and call for models that can perceive contextually and culturally beyond the surface of what they hear and see. Project page: avmemeexam.github.io/public

多模态文化理解音视频基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。