首个跨语言多模态指令遵循基准,用于评估科学演讲中的模型理解能力。
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- 基于科学演讲构建跨语言多模态指令数据集
- 覆盖4语言3模态4类任务,支持长短输入统一评估
- 含人工标注,适合研究多模态大模型的跨语言性能
大型语言模型的发展为多模态大模型(MLLMs)奠定了基础,将文本、语音和视觉统一于单一框架中。随着模型向跨任务、复杂场景的通用指令遵循演进,评估其跨语言与多模态能力成为关键挑战。现有基准存在局限:多限于英语,单模态为主,以短输入为主,缺乏人工标注,难以全面评估模型在语言、模态与任务复杂度上的表现。为此,我们提出MCIF(Multimodal Crosslingual Instruction-Following),首个基于自然语言处理等领域的科学演讲构建的跨语言人工标注基准。MCIF在跨语言、多模态场景下评估指令遵循能力,涵盖不同输入长度,包含四类宏观任务:识别、翻译、问答与摘要。覆盖三种核心模态(语音、视觉、文本)与四种语言(英语、德语、意大利语、中文),全维度对齐。该并行设计可系统评估MLLMs跨语言指令理解与多模态信息融合能力。对23个模型的基准测试与分析揭示了跨模态与任务的普遍挑战,表明未来模型仍有巨大提升空间。MCIF以CC-BY 4.0许可证开放,推动开源研究。
原文摘要 · Abstract (English)
Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework. As these models are rapidly evolving toward general-purpose instruction following across diverse and complex tasks, a key frontier is evaluating their crosslingual and multimodal capabilities over both short- and long-form inputs. However, existing benchmarks fall short in evaluating these dimensions jointly: they are often limited to English, mostly focus on a single modality at a time, rely on short-form inputs, or lack human annotations--hindering comprehensive assessment of model performance across languages, modalities, and task complexity. To address these gaps, we introduce MCIF (Multimodal Crosslingual Instruction Following), the first crosslingual human-annotated benchmark based on scientific talks on NLP and beyond. MCIF evaluates instruction following in crosslingual, multimodal settings over different input lengths and spans four macro-tasks: recognition, translation, question answering, and summarization. It covers three core modalities (speech, vision, and text) and four diverse languages (English, German, Italian, and Chinese), fully aligned across all dimensions. This parallel design enables a systematic evaluation of MLLMs' abilities to interpret instructions across languages and effectively integrate multimodal contextual information. Our benchmarking and analysis of 23 models highlight universal challenges across modalities and tasks, indicating substantial room for improvement in future MLLMs development. MCIF is released under CC-BY 4.0 license to promote open research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。