arXiv:2606.15144cs.CLcs.AI2026-06

评测大模型对菲律宾语构词结构的理解能力,发现主流模型仍难突破形态组合瓶颈。

PACUTE: Phonology-, Affix-, and Character-level Understanding of Tokens for Filipino

  • 构建4600任务的诊断基准,分六层评估音位、词缀、字符级理解
  • 开源模型在词素分解上仅接近随机水平,商用模型虽能识别词缀但组合能力弱
  • 揭示菲律宾语中词形变化和音节化是当前模型的核心短板,适合语言学与NLP交叉研究者

大语言模型将文本处理为子词标记序列,常掩盖字符级与形态结构,尤其在非连缀性语法语言中,标准分词器会系统性错位词素边界。本文提出PACUTE,一个包含4,600个任务的诊断基准,用于评估菲律宾语的形态理解能力。菲律宾语具有活跃的中缀、重叠与变音符号驱动的词汇差异,通常未体现在书面文本中。PACUTE采用六层级的层次化诊断框架,定位形态理解失效的具体环节。评估开源权重模型与前沿商业模型发现:开源模型在词素分解上表现接近随机水平,无论规模大小;前沿模型表现显著更好,常能在包含匹配评分下恢复单个词缀,但在词素变换与音节化等组合任务上仍远低于字符级理论上限。结果表明,词形的生产性组合而非单纯字符访问,仍是菲律宾语词结构理解的持续瓶颈。

原文摘要 · Abstract (English)

Large language models (LLMs) process text as sequences of subword tokens, which can obscure the character-level and morphological structure that underlies word formation. This limitation is most acute for languages with non-concatenative morphology, where standard tokenizers systematically misalign token boundaries with morpheme boundaries. We introduce PACUTE, a diagnostic benchmark of 4,600 tasks designed to evaluate morphological understanding in Filipino, a language characterized by productive infixation, reduplication, and diacritic-driven lexical distinctions that are typically absent from written text. PACUTE includes a hierarchical diagnostic framework of six compositional levels that localizes where morphological understanding breaks down. Evaluating open-weight LLMs and frontier commercial models, we find that open-weight models perform near chance on morpheme decomposition regardless of scale. Frontier models perform much better, often recovering individual affixes under contains-match scoring, but remain far below their character-level ceilings on compositional tasks of morpheme transformations and syllabification. These results identify productive morphological composition, rather than character access alone, as the persistent bottleneck for Filipino word-structure understanding.

形态学多语言语言模型菲律宾语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。