arXiv:2511.15552cs.CLcs.AI2025-11Conference of the …被引 4

首个俄语多模态评测框架,填补语言空白。

Multimodal Evaluation of Russian-language Architectures

  • 构建18个全新任务,覆盖文本、图像、音频、视频四模态。
  • 针对俄语文化特性设计数据集,统一评测标准与指标。
  • 适合研究多模态模型在斯拉夫语系中的表现与风险。

多模态大语言模型(MLLMs)正迅速发展,但其智能水平、局限性与潜在风险仍不明确。针对俄语领域尚无多模态评测基准的问题,我们提出MERA Multi,一个面向俄语架构的开源多模态评估框架。该基准基于指令,涵盖文本、图像、音频和视频四种模态,包含18项全新构建的评估任务,适用于通用模型及图像到文本、视频到文本、音频到文本等特定模态架构。贡献包括:(i) 提出通用的多模态能力分类体系;(ii) 从零构建18个数据集,注重俄语文化和语言特性,统一提示与度量标准;(iii) 提供闭源与开源模型的基线结果;(iv) 设计防泄露方法,包括对私有数据集进行水印处理。尽管当前聚焦俄语,该框架为构建类型多样的语言(尤其是斯拉夫语族)多模态评测提供可复现的方法论。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are currently at the center of research attention, showing rapid progress in scale and capabilities, yet their intelligence, limitations, and risks remain insufficiently understood. To address these issues, particularly in the context of the Russian language, where no multimodal benchmarks currently exist, we introduce MERA Multi, an open multimodal evaluation framework for Russian-spoken architectures. The benchmark is instruction-based and encompasses default text, image, audio, and video modalities, comprising 18 newly constructed evaluation tasks for both general-purpose models and modality-specific architectures (imageto-text, video-to-text, and audio-to-text). Our contributions include: (i) a universal taxonomy of multimodal abilities; (ii) 18 datasets created entirely from scratch with attention to Russian cultural and linguistic specificity, unified prompts, and metrics; (iii) baseline results for both closed-source and open-source models; (iv) a methodology for preventing benchmark leakage, including watermarking for private sets. While our current focus is on Russian, the proposed benchmark provides a replicable methodology for constructing multimodal benchmarks in typologically diverse languages, particularly within the Slavic language family.

多模态俄语评测基准语言多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。