arXiv:2608.21853cs.CL2026-08

构建波兰文化多模态理解基准,揭示模型在本地化场景下的短板

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

论文配图:PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding
图 1 · 摘自论文原文
  • 设计900个手工任务,覆盖波兰文化语境下的图文音多模态理解
  • 顶尖商用模型在视觉问答中表现良好,但复杂音频与文档理解仍存巨大差距
  • 开源评估框架,助力本地化多模态AI研究

大型语言模型正从文本处理拓展至图像、音频等多模态能力。尽管文本理解和生成已广泛研究,但在非英语语言与文化背景下的多模态数据处理尚未得到全面评估。本文提出PUMA(波兰统一多模态评估),一个包含900个精心设计任务的新基准,用于探测多模态模型在波兰文化与语言背景下的能力极限。该数据集评估了文化理解力及对文本、图像、音频和视觉富文档的处理能力。我们对前沿商业模型、开源权重模型及专用小型系统进行了广泛评估,发现显著性能差距:虽然顶级商用模型在视觉问答中得分较高,但多数模型在复杂音频或文档理解上表现不佳。我们开源了评估框架,以推动本地化多模态AI研究。

原文摘要 · Abstract (English)

Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understanding and generation have been extensively studied, multimodal data processing capabilities, particularly in the context of cultures and languages other than English, have not yet been evaluated comprehensively. In this paper, we propose PUMA (Polish Unified Multimodal Assessment), a novel benchmark of 900 hand-crafted tasks designed to probe the limits of multimodal models in the Polish cultural and linguistic context. The dataset evaluates both cultural understanding and practical skill in processing text, images, audio, and visually rich documents. Our extensive evaluation of frontier commercial models, open-weights models, and specialized smaller systems highlights a significant performance gap. While top commercial models achieve high scores in visual question answering, most models struggle with complex audio or document understanding. We open-source our evaluation framework to advance localized multimodal AI research.

多模态文化理解波兰语评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。