评测大模型对文化语境中隐喻语言的实用处理能力,发现其理解与使用存在显著差距。
Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language
- 设计多语言隐喻任务,评估模型对阿拉伯语和英语习语的理解与实际运用。
- 埃及方言习语准确率比阿拉伯谚语低10.28%,实际使用任务准确率比理解低14.07%。
- 发布首个埃及阿拉伯语习语数据集Kinayat,支持未来文化推理研究。
我们全面评估了大语言模型(LLMs)在处理基于文化的语言方面的能力,重点关注其对蕴含本地知识与文化细微差别的隐喻表达的理解与实际运用能力。以隐喻语言为文化细微差别的代理指标,我们在阿拉伯语和英语中设计了上下文理解、语用使用和语义内涵解读三类评估任务。评估了22个开源与闭源模型在埃及阿拉伯语习语、多方言阿拉伯谚语和英语谚语上的表现。结果表明存在一致层级:阿拉伯谚语平均准确率比英语谚语低4.29%,埃及习语准确率比阿拉伯谚语低10.28%。在语用使用任务中,准确率相比理解任务下降14.07%,但提供语境句可使准确率提升10.66%。模型在语义内涵理解上也表现不佳,与人类标注者最高仅达85.58%一致率,而该任务的人类标注一致性为100%。这些发现表明,隐喻语言是检测文化推理能力的有效诊断工具:尽管模型常能解释隐喻含义,但在恰当使用上仍面临挑战。为支持后续研究,我们发布了Kinayat——首个专为埃及阿拉伯语习语的隐喻理解与语用使用评估设计的数据集。
原文摘要 · Abstract (English)
We present a comprehensive evaluation of the ability of large language models (LLMs) to process culturally grounded language, specifically to understand and pragmatically use figurative expressions that encode local knowledge and cultural nuance. Using figurative language as a proxy for cultural nuance and local knowledge, we design evaluation tasks for contextual understanding, pragmatic use, and connotation interpretation in Arabic and English. We evaluate 22 open- and closed-source LLMs on Egyptian Arabic idioms, multidialectal Arabic proverbs, and English proverbs. Our results show a consistent hierarchy: the average accuracy for Arabic proverbs is 4.29% lower than for English proverbs, and performance for Egyptian idioms is 10.28% lower than for Arabic proverbs. For the pragmatic use task, accuracy drops by 14.07% relative to understanding, though providing contextual idiomatic sentences improves accuracy by 10.66%. Models also struggle with connotative meaning, reaching at most 85.58% agreement with human annotators on idioms with 100% inter-annotator agreement. These findings demonstrate that figurative language serves as an effective diagnostic for cultural reasoning: while LLMs can often interpret figurative meaning, they face challenges in using it appropriately. To support future research, we release Kinayat, the first dataset of Egyptian Arabic idioms designed for both figurative understanding and pragmatic use evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。