arXiv:2409.16646cs.CL2024-09NAACL被引 5

跨语言图像描述揭示文化对视觉认知的影响

Cross-Lingual and Cross-Cultural Variation in Image Descriptions

  • 构建多语言图像描述分析方法,精准识别图片中实体与描述匹配度
  • 发现地理或语言亲缘关系近的语言更常提及相同实体,如人比服饰更普遍
  • 验证了基本类别理论和环境影响感知的旧有假设,适合语言学与认知研究者

不同语言使用者对所见之物的描述是否不同?以往行为与认知研究虽报告文化对感知的影响,但规模小且难复现。本文首次开展大规模跨语言图像描述实证研究,使用包含31种语言和多元地域图像的多模态数据集,开发方法以准确识别描述中提及且存在于图像中的实体,并分析其跨语言差异。结果显示,地理或语言上相近的语言更倾向于提及相同实体。我们还发现某些实体类别具有普遍高显著性(如生物体)、低显著性(如服饰配件)或高跨语言变异性(如风景)。案例研究显示,日语比英语更频繁提及衣物。该方法验证了早期小规模研究:1)罗施等(1976)的基本类别理论,即偏好既不泛化也不过度具体实体;2)宫本等(2006)的环境塑造感知模式假说,如实体数量差异。总体表明,实体提及同时存在普遍规律与文化特异性。

原文摘要 · Abstract (English)

Do speakers of different languages talk differently about what they see? Behavioural and cognitive studies report cultural effects on perception; however, these are mostly limited in scope and hard to replicate. In this work, we conduct the first large-scale empirical study of cross-lingual variation in image descriptions. Using a multimodal dataset with 31 languages and images from diverse locations, we develop a method to accurately identify entities mentioned in captions and present in the images, then measure how they vary across languages. Our analysis reveals that pairs of languages that are geographically or genetically closer tend to mention the same entities more frequently. We also identify entity categories whose saliency is universally high (such as animate beings), low (clothing accessories) or displaying high variance across languages (landscape). In a case study, we measure the differences in a specific language pair (e.g., Japanese mentions clothing far more frequently than English). Furthermore, our method corroborates previous small-scale studies, including 1) Rosch et al. (1976)'s theory of basic-level categories, demonstrating a preference for entities that are neither too generic nor too specific, and 2) Miyamoto et al. (2006)'s hypothesis that environments afford patterns of perception, such as entity counts. Overall, our work reveals the presence of both universal and culture-specific patterns in entity mentions.

跨语言视觉认知文化差异图像描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。