arXiv:2510.06371cs.CLcs.AI2025-10被引 2

构建多语言多模态文化常识问答数据集,提升模型跨文化理解能力。

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

  • 基于半自动框架生成跨语言语音图文问答数据
  • 含1480万组问答,覆盖3.7万条语音问题和20小时克隆语音
  • 聚焦文化常识与日常推理,适合评估跨语言智能系统

大规模多模态模型在视觉问答任务中表现强劲,但在需要文化背景、视觉信息与日常知识的查询上仍受限,尤其在低资源和代表性不足的语言中。我们提出OASIS,一个大规模、文化相关的多模态问答数据集,包含图像、文本和语音。OASIS基于EverydayMMQA框架构建,采用多阶段人机协作验证,实现可扩展的本地化语音与视觉问答资源生成。数据集包含约92万张真实图像和1480万组问答对,其中370万条为语音问题,涵盖383小时真人录音与20,000小时语音克隆数据,来自42名说话者。支持文本仅、语音仅、文本+图像、语音+图像四种输入形式。覆盖英语和阿拉伯语变体,横跨18个国家,包括现代标准阿拉伯语(MSA)及方言。旨在评估超越物体识别的实用、常识性与文化相关推理能力。我们在OASIS上测试了四款闭源模型、三款开源模型及一款微调模型。框架与数据集将向社区公开。https://huggingface.co/datasets/QCRI/OASIS

原文摘要 · Abstract (English)

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low-resource and underrepresented languages. We introduce OASIS, a large-scale culturally grounded multimodal QA dataset covering images, text, and speech. OASIS is built with EverydayMMQA, a scalable semi-automatic framework for creating localized spoken and visual QA resources, supported by multi-stage human-in-the-loop validation. OASIS contains approximately 0.92M real images and 14.8M QA pairs, including 3.7M spoken questions, with 383 hours of human-recorded speech, and 20K hours of voice-cloned speech, from 42 speakers. It supports four input settings: text-only, speech-only, text+image, and speech+image. The dataset focuses on English and Arabic varieties across 18 countries, covering Modern Standard Arabic (MSA) as well as dialectal Arabic. It is designed to evaluate models beyond object recognition, targeting pragmatic, commonsense, and culturally grounded reasoning in real-world scenarios. We benchmark four closed-source models, three open-source models, and one fine-tuned model on OASIS. The framework and dataset will be made publicly available to the community. https://huggingface.co/datasets/QCRI/OASIS

多模态文化常识语音问答低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。