测试多语言大模型在文化类流程文本上的理解能力,发现其表现受限于语言资源和文化背景。
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
- 构建CAPTex基准,评估多语言模型对跨文化流程文本的处理能力。
- 低资源语言中模型性能显著下降,不同文化领域表现差异明显。
- 对话式多选题比直接提问更易让模型正确回答,凸显任务设计影响。
尽管多语言大语言模型(mLLMs)在多种自然语言处理任务中表现出色,但其对流程类文本——特别是包含文化特有内容的文本——的理解能力仍缺乏深入研究。这类文本如仪式、传统工艺和社交礼仪,需依赖深层文化背景理解,对mLLMs构成重大挑战。本文提出CAPTex基准,通过多种方法评估mLLMs在多语言环境下对文化多样性流程文本的处理与推理能力。结果表明:(1) mLLMs在文化语境化流程文本上表现不佳,尤其在低资源语言中性能明显下降;(2) 模型在不同文化领域间表现波动,某些领域更具挑战性;(3) 在对话框架下的多选题中,模型表现优于直接提问任务。这些发现揭示了mLLMs在处理文化细微流程文本时的当前局限,并强调了像CAPTex这样的文化感知基准对提升其跨语言、跨文化适应性的必要性。
原文摘要 · Abstract (English)
Despite the impressive performance of multilingual large language models (mLLMs) in various natural language processing tasks, their ability to understand procedural texts, particularly those with culture-specific content, remains largely unexplored. Texts describing cultural procedures, including rituals, traditional craftsmanship, and social etiquette, require an inherent understanding of cultural context, presenting a significant challenge for mLLMs. In this work, we introduce CAPTex, a benchmark designed to evaluate mLLMs' ability to process and reason about culturally diverse procedural texts across multiple languages using various methodologies to assess their performance. Our findings indicate that (1) mLLMs face difficulties with culturally contextualized procedural texts, showing notable performance declines in low-resource languages, (2) model performance fluctuates across cultural domains, with some areas presenting greater difficulties, and (3) language models exhibit better performance on multiple-choice tasks within conversational frameworks compared to direct questioning. These results underscore the current limitations of mLLMs in handling culturally nuanced procedural texts and highlight the need for culturally aware benchmarks like CAPTex to enhance their adaptability and comprehension across diverse linguistic and cultural landscapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。