arXiv:2504.05995cs.CLcs.AI2025-04

构建多模态本地知识问答数据集,提升大模型在不同文化语境下的表现

NativQA Framework: Enabling LLMs and VLMs with Native, Local, and Everyday Knowledge

  • 基于用户种子问题,通过搜索引擎采集本地化多模态信息
  • 覆盖39个地区、24个国家、7种语言,收集超30万文本对、31.2万张图片和2.9万段音视频
  • 开源框架支持文化适配训练,适合做多语言与跨文化研究的团队使用

大语言模型的快速发展引发了对文化偏见、公平性及多语言、欠资源地区性能的担忧。填补这些空白需要大规模、扎根于多语言、本地和文化背景的资源。我们系统化并扩展了早期的NativQA框架,加入图像、音频和视频支持,实现以母语为基础、跨文化地域对齐的问答数据集规模化构建。给定用户定义的种子问题,该框架利用搜索引擎收集特定地点的日常信息。我们在24个国家的39个地点、7种语言下进行评估,涵盖从极低资源到高资源场景,共收集约30万条文本问答对、约31.2万张图像以及约2.9万段带音频的视频。所生成资源可用于大模型基准测试与微调。该框架已向社区公开(https://gitlab.com/nativqa/nativqa-framework),演示视频可访问:https://shorturl.at/DAVn9。

原文摘要 · Abstract (English)

The rapid progress of large language models (LLMs) raises concerns about cultural bias, fairness, and performance in diverse languages and underrepresented regions. Addressing these gaps requires large-scale resources grounded in multilingual, local, and cultural contexts. We systematize and extend the earlier NativQA framework to multimodality by adding image, audio, and video support, enabling scalable construction of culturally and regionally aligned QA datasets in native languages. Given user-defined seed queries, the framework uses search engines to collect location-specific everyday information. We evaluate it across 39 locations in 24 countries and 7 languages, spanning extremely low-resource to high-resource settings, and collect over $\sim$300K text QA pairs, $\sim$312K images, and $\sim$29K videos with associated audio. The developed resources can be used for LLMs benchmarking and further fine-tuning. The framework has been made publicly available for the community (https://gitlab.com/nativqa/nativqa-framework). Demo video is available here: \href{https://shorturl.at/DAVn9}{https://shorturl.at/DAVn9}.

多模态本地知识跨文化数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。