构建东南亚多文化视觉语言数据集,提升AI对当地文化的理解
Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia
- 结合众包、爬取与生成,探索适合东南亚文化的图像采集方法
- 采集128万张文化相关图像,规模超现有数据集50倍以上
- 发现爬取法文化相关性达85%且更高效,生成图像仍难准确反映文化细节
东南亚地区语言与文化多样性突出,但在视觉语言(VL)研究中仍严重缺位,导致人工智能模型难以捕捉当地文化特征。为此,我们提出SEA-VL,一个开源的高质量、文化相关的东南亚多语言视觉语言数据集。通过来自东南亚国家的贡献者参与,确保数据的文化相关性与多样性,推动未充分代表语言在VL研究中的包容性。除了众包外,还探索了自动采集方法:图像爬取可实现约85%的文化相关性,且成本和时间效率更高;尽管生成模型进展显著,合成图像仍无法准确反映东南亚的细微传统与文化背景。最终,我们共收集128万张具有文化相关性的东南亚图像,规模超过现有数据集50倍。SEA-VL旨在弥合东南亚地区的代表性差距,推动更真实、包容的AI系统发展。
原文摘要 · Abstract (English)
Southeast Asia (SEA) is a region of extraordinary linguistic and cultural diversity, yet it remains significantly underrepresented in vision-language (VL) research. This often results in artificial intelligence (AI) models that fail to capture SEA cultural nuances. To fill this gap, we present SEA-VL, an open-source initiative dedicated to developing high-quality, culturally relevant data for SEA languages. By involving contributors from SEA countries, SEA-VL aims to ensure better cultural relevance and diversity, fostering greater inclusivity of underrepresented languages in VL research. Beyond crowdsourcing, our initiative goes one step further in the exploration of the automatic collection of culturally relevant images through crawling and image generation. First, we find that image crawling achieves approximately ~85% cultural relevance while being more cost- and time-efficient than crowdsourcing. Second, despite the substantial progress in generative vision models, synthetic images remain unreliable in accurately reflecting SEA cultures. The generated images often fail to reflect the nuanced traditions and cultural contexts of the region. Collectively, we gather 1.28M SEA culturally-relevant images, more than 50 times larger than other existing datasets. Through SEA-VL, we aim to bridge the representation gap in SEA, fostering the development of more inclusive AI systems that authentically represent diverse cultures across SEA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。