arXiv:2509.08960cs.CL2025-09被引 3

用巴西谚语测评大模型对本地语言的理解能力

BRoverbs -- Measuring how much LLMs understand Portuguese proverbs

  • 构建巴西谚语数据集,评估大模型对本土表达的理解
  • 填补葡萄牙语本地化评估空白,突出文化与语言细节
  • 适合关注多语言、跨文化大模型评测的研究者

大型语言模型在不同语言和文化背景下的表现差异显著,凸显了建立成熟评估框架的必要性。针对葡萄牙语,现有评估仍局限于翻译数据集,难以捕捉语言细微差别与文化内涵。而本地数据集多聚焦于标准化考试或社交媒体情感分析,缺乏对更广泛语言理解能力的考察。为此,本文提出BRoverbs——一个专用于评估大模型对巴西谚语理解能力的数据集。谚语蕴含丰富的文化智慧、隐喻表达与复杂句法结构,能有效检验模型对区域性语言现象的理解深度。该基准测试已公开于Hugging Face,旨在推动面向区域语境的基准评测发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit significant performance variations depending on the linguistic and cultural context in which they are applied. This disparity signals the necessity of mature evaluation frameworks that can assess their capabilities in specific regional settings. In the case of Portuguese, existing evaluations remain limited, often relying on translated datasets that may not fully capture linguistic nuances or cultural references. Meanwhile, native Portuguese-language datasets predominantly focus on structured national exams or sentiment analysis of social media interactions, leaving gaps in evaluating broader linguistic understanding. To address this limitation, we introduce BRoverbs, a dataset specifically designed to assess LLM performance through Brazilian proverbs. Proverbs serve as a rich linguistic resource, encapsulating cultural wisdom, figurative expressions, and complex syntactic structures that challenge the model comprehension of regional expressions. BRoverbs aims to provide a new evaluation tool for Portuguese-language LLMs, contributing to advancing regionally informed benchmarking. The benchmark is available at https://huggingface.co/datasets/Tropic-AI/BRoverbs.

语言理解大模型评测葡萄牙语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。