arXiv:2505.17784cs.CL2025-05ACL被引 4

构建多语言分词理解基准,发现大模型在不同语言中问题类型各异。

EXECUTE: A Multilingual Benchmark for LLM Token Understanding

  • 扩展英文字符理解测试至多语言,支持任意语言快速接入
  • 发现非英语语言中存在词级处理问题,部分语言无明显缺陷
  • 首次系统评估中日韩文字部件理解能力,揭示深层认知差异

CUTE基准显示大语言模型在英语字符理解上表现不佳。本文将其扩展至多种语言,涵盖多样文字与书写系统,提出新基准EXECUTE。简化框架使该测试可轻松拓展至任何语言。对多个大模型的测试表明,其他语言的问题并非总在字符层面——部分语言存在词级处理缺陷,部分则无显著问题。此外,针对中文、日文和韩文的子字符任务分析,评估了模型对汉字部件的理解能力。

原文摘要 · Abstract (English)

The CUTE benchmark showed that LLMs struggle with character understanding in English. We extend it to more languages with diverse scripts and writing systems, introducing EXECUTE. Our simplified framework allows easy expansion to any language. Tests across multiple LLMs reveal that challenges in other languages are not always on the character level as in English. Some languages show word-level processing issues, some show no issues at all. We also examine sub-character tasks in Chinese, Japanese, and Korean to assess LLMs' understanding of character components.

语言理解多语言分词评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。