首个面向视频理解的意大利语多模态推理基准,揭示模型在文化语境下的脆弱性
All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark
- 构建意大利文化视频数据集,评估视觉语言模型的跨模态推理能力
- 涵盖12类推理任务,通过双任务联合评估验证与生成的一致性
- 专为本土化语境设计,适合研究多模态模型在真实场景中的泛化能力
我们提出MAIA(多模态人工智能评估),一个原生意大利语的视频多模态推理基准,用于细粒度评估视觉语言模型在视频理解中的推理能力。与现有视频基准不同,MAIA在设计、推理类别、评估指标及视频的语言文化背景上具有独特性。它在相同视频问答数据集上评估视觉语言模型的两项对齐任务:视觉陈述验证和开放式视觉问答。该基准包含十二类推理类别,旨在通过突出视觉输入的作用,解耦语言与视觉的关系。得益于精心设计,它通过聚合指标同时评估模型的一致性和视觉引导的自然语言理解与生成能力,揭示出模型表现不佳的薄弱环节。最后,视频素材经严格筛选以反映意大利文化特征,语言数据由母语者生成。
原文摘要 · Abstract (English)
We introduce MAIA (Multimodal AI Assessment), a native-Italian benchmark designed for fine-grained investigation of the reasoning abilities of visual language models on videos. MAIA differs from other available video benchmarks for its design, its reasoning categories, the metric it uses, and the language and culture of the videos. MAIA evaluates Vision Language Models (VLMs) on two aligned tasks: a visual statement verification task, and an open-ended visual question-answering task, both on the same set of video-related questions. It considers twelve reasoning categories that aim to disentangle language and vision relations by highlighting the role of the visual input. Thanks to its carefully taught design, it evaluates VLMs' consistency and visually grounded natural language comprehension and generation simultaneously through an aggregated metric revealing low results that highlight models' fragility. Last but not least, the video collection has been carefully selected to reflect the Italian culture, and the language data are produced by native-speakers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。