评测大模型对中文发音的理解能力,发现其表现远不如人类。
Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese

- 构建中文发音理解专项评测集,覆盖同音、押韵、发音相似三维度。
- 模型能准确复述发音但难以灵活运用语音知识,与人类差距明显。
- 揭示模型语音感知机制缺陷,为后续研究指明方向。
语言是思想的载体,与声音、符号和意义紧密相连。然而,当前大型语言模型(LLM)研究主要关注语义和拼写,忽视了声音层面。现有评测要么可靠死记硬背完成,要么与其他能力混杂,无法真实衡量模型的语音理解能力。为此,我们提出Phun-Bench,一个专为中文设计的基准测试,涵盖同音、押韵和发音相似三个维度的多样化任务,系统评估模型的语音理解水平。结果表明,尽管模型在正确复述发音方面表现良好,但在灵活、直观地运用语音知识方面仍显著逊色于人类。通过深入分析,我们提出了关于模型语音理解与‘感知’机制的假设,揭示了一个亟待探索的研究前沿。
原文摘要 · Abstract (English)
Language is a vehicle for thought, intricately tied to sounds, symbols, and meaning. However, most large language model (LLM) research focuses on meaning (semantics) and symbols (spelling) while largely overlooking sounds. Existing benchmarks on LLMs' phonological abilities are either solvable through rote memorization or intertwined with other abilities, making them inadequate to measure LLMs' genuine ability in phonological understanding. Here, we present Phun-Bench, a purpose-built Chinese benchmark with diverse tasks and settings across three dimensions (Homophony, Rhyme, and Phonetic Similarity), designed to systematically evaluate LLMs' phonological understanding. Our results show that while LLMs excel at recalling correct pronunciations, they generally struggle to leverage phonological knowledge in the flexible and intuitive way that human speakers do. Moreover, through detailed analyses, we propose a hypothesis regarding the underlying mechanism of LLMs' phonological understanding and "perception", highlighting an underexplored frontier for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。