构建首个多语言视频理解基准,覆盖14种语言与多元文化场景。
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
- 构建跨14种语言的视频多模态评测基准ViMUL-Bench,含8000个经母语者验证的样本。
- 提出ViMUL模型,在高/低资源语言间实现更好的视频理解平衡性能。
- 数据集涵盖15类文化主题,支持长/短视频与开放问答,推动多语言包容性研究。
大型多模态模型(LMMs)在视觉内容理解与生成方面表现优异,但多数仍局限于英语。尽管少数工作探索了多语言图像LMM,但视频领域的跨语言与跨文化包容性尚未被系统研究。为此,我们提出首个多语言视频LMM基准ViMUL-Bench,覆盖14种语言(包括英语、中文、西班牙语、法语、德语、印地语、阿拉伯语、俄语、孟加拉语、乌尔都语、僧伽罗语、泰米尔语、瑞典语、日语),涵盖15个类别,包含生活方式、节庆、食物、仪式、地标与文化名人等8类文化多样性主题。该基准包含8000个经母语者人工验证的样本,支持短/长格式开放问题与多项选择题,覆盖短、中、长视频时长。同时,我们构建了包含120万样本的机器翻译多语言视频训练集,并开发简单高效的多语言视频LMM模型ViMUL,其在高/低资源语言间表现出更优的视频理解权衡。我们期望该基准、模型与数据集能推动未来更具文化与语言包容性的多语言视频模型研究。相关资源将公开发布于https://mbzuai-oryx.github.io/ViMUL/。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in English language. While few recent works explore multilingual image LMMs, to the best of our knowledge, moving beyond the English language for cultural and linguistic inclusivity is yet to be investigated in the context of video LMMs. In pursuit of more inclusive video LMMs, we introduce a multilingual Video LMM benchmark, named ViMUL-Bench, to evaluate Video LMMs across 14 languages, including both low- and high-resource languages: English, Chinese, Spanish, French, German, Hindi, Arabic, Russian, Bengali, Urdu, Sinhala, Tamil, Swedish, and Japanese. Our ViMUL-Bench is designed to rigorously test video LMMs across 15 categories including eight culturally diverse categories, ranging from lifestyles and festivals to foods and rituals and from local landmarks to prominent cultural personalities. ViMUL-Bench comprises both open-ended (short and long-form) and multiple-choice questions spanning various video durations (short, medium, and long) with 8k samples that are manually verified by native language speakers. In addition, we also introduce a machine translated multilingual video training set comprising 1.2 million samples and develop a simple multilingual video LMM, named ViMUL, that is shown to provide a better tradeoff between high-and low-resource languages for video understanding. We hope our ViMUL-Bench and multilingual video LMM along with a large-scale multilingual video training set will help ease future research in developing cultural and linguistic inclusive multilingual video LMMs. Our proposed benchmark, video LMM and training data will be publicly released at https://mbzuai-oryx.github.io/ViMUL/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。