arXiv:2607.00664cs.CL2026-07

评测大模型对日语汉字读音与发音理解能力,发现主流模型表现不佳。

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

论文配图:YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese
图 1 · 摘自论文原文
  • 设计四类任务,专门测试日语汉字读音理解能力。
  • 多语言和日本专用模型在读音任务中表现均低于预期。
  • 商业模型在需考虑汉字读音的生成任务中同样表现差。

我们提出 YOMI-Bench,一个用于评估大语言模型(LLMs)在日语中汉字读音与语音理解能力的基准。日语中单个汉字常有多种可能读音,仅从表层文本难以准确推断正确读音。由于这一语言特性,已有实证表明大模型在日语汉字读音任务上表现较差。所提出的 YOMI-Bench 包含四项专为评估日语汉字读音能力而设计的任务。我们在该基准上评估了一款多语言开源模型、四款日本专用开源模型以及五款商业大模型。结果显示,即使日本专用模型性能也普遍较低,商业模型在需要考虑汉字读音的生成任务中同样表现不佳。

原文摘要 · Abstract (English)

We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese, a single kanji character often has multiple possible readings, making it difficult to infer the correct reading from surface-level text alone. Due to these linguistic characteristics, it is empirically known that LLMs exhibit low performance in kanji reading for Japanese. The proposed YOMI-Bench consists of four tasks specifically designed to evaluate kanji reading performance in Japanese. In our evaluation using YOMI-Bench, we assessed one multilingual open LLM, four Japanese-specific open LLMs, and five commercial LLMs. As a result, we found that even Japanese-specific models show low performance, and that commercial models also perform poorly on generation tasks that require consideration of kanji readings.

日语理解汉字读音大模型评测语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。