研究大模型如何像人一样理解错位的单词,发现其依赖固定字形特征。
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia
- 提出新指标SemRecScore量化语义重建程度
- 字形信息是大模型理解错位词的核心因素
- 大模型用特定注意力头处理字形,稳定且固定
人类读者能高效理解拼写错乱的单词(即拼写糖现象),主要依赖词形信息;若仅靠词形不足,则进一步借助上下文线索。尽管先进大语言模型(LLMs)具备类似能力,其内在机制仍不清晰。为此,我们设计受控实验,分析词形与上下文信息在语义重建中的作用,并考察大模型注意力模式。首先提出SemRecScore,一种可靠量化语义重建程度的指标,并验证其有效性。基于该指标,研究发现词形是大模型语义重建的核心因素。进一步分析表明,大模型通过特定注意力头提取并处理词形信息,该机制在不同单词错位程度下保持稳定。这与人类读者在词形与上下文间动态权衡的适应性策略形成对比,提示可通过引入类人、上下文感知机制提升大模型性能。
原文摘要 · Abstract (English)
Human readers can efficiently comprehend scrambled words, a phenomenon known as Typoglycemia, primarily by relying on word form; if word form alone is insufficient, they further utilize contextual cues for interpretation. While advanced large language models (LLMs) exhibit similar abilities, the underlying mechanisms remain unclear. To investigate this, we conduct controlled experiments to analyze the roles of word form and contextual information in semantic reconstruction and examine LLM attention patterns. Specifically, we first propose SemRecScore, a reliable metric to quantify the degree of semantic reconstruction, and validate its effectiveness. Using this metric, we study how word form and contextual information influence LLMs' semantic reconstruction ability, identifying word form as the core factor in this process. Furthermore, we analyze how LLMs utilize word form and find that they rely on specialized attention heads to extract and process word form information, with this mechanism remaining stable across varying levels of word scrambling. This distinction between LLMs' fixed attention patterns primarily focused on word form and human readers' adaptive strategy in balancing word form and contextual information provides insights into enhancing LLM performance by incorporating human-like, context-aware mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。