arXiv:2609.01788cs.CLcs.AI2026-09

首个针对印地语系语言的语用能力评测,揭示大模型在文化语境理解上的普遍短板。

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

论文配图:VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages
图 1 · 摘自论文原文
  • 构建跨四种印地语系语言的语用诊断基准,涵盖指涉、言语行为等五类现象。
  • 多语言大模型在指涉和暗示等语用任务上准确率普遍低于40%,体现文化理解缺陷。
  • 适合研究多语言AI、跨文化语用与本土化评测的学者与工程师参考。

真实交流常需语用推理:理解基于语境和文化惯例的隐含意义,而非字面意思。现有语用评估主要局限于英语和高资源语言,忽视了印地语系语言的丰富性。我们提出VakyArth,首个面向印地语系语言的语用评测基准,覆盖印地语、旁遮普语、泰米尔语和马拉雅拉姆语。该基准通过选择题、自然语言推理和翻译任务,评估模型在指涉、言语行为、暗示、社会语用和连贯性五个方面的能力,所有题目均由母语者编写。在不同家族与规模的多语言大模型中,我们发现其在根植于印度语言文化惯例的语用理解上存在系统性失败。分析显示:选择题准确率普遍高于自然语言推理;翻译表现无法可靠反映语用理解水平;印欧语系语言模型在翻译上优于德拉维达语系语言。此外,自动翻译指标会忽略流畅但语用不忠实的输出,尤其在暗示和指涉任务中表现不佳。

原文摘要 · Abstract (English)

Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cultural diversity. We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam. VakyArth evaluates models across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence; through multiple-choice questions, natural language inference, and translation, with all items authored by native speakers. Across multilingual large language models (LLMs) of varying families and sizes, we find consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions. Our analysis shows systematic differences across languages and tasks: MCQ accuracy exceeds NLI accuracy in all model-language combinations, translation performance does not reliably track pragmatic understanding, and Indo-Aryan languages show a translation advantage over Dravidian languages. We further show that automatic translation metrics can miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.

语用学多语言大模型评测印地语系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。