研究动词+up短语的存储机制,发现频率与可预测性决定其表征方式。
The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models

- 通过分析文本与语音模型内部表示,探究动词+up短语的整合理解机制。
- 高频且可预测的动词+up短语在模型中形成独立表征,支持使用型语言理论。
- 适用于语言认知、自然语言处理及语音识别领域的研究人员。
语言处理的核心能力之一是在存储表征与抽象知识之间进行权衡:既要能检索已有表征,又要能通过生成规则创造新表达。尽管近期研究关注语言模型中的抽象知识,对整体性存储的关注仍不足。本文探测了文本型大语言模型与语音识别模型的内部表征,检验动词+up短语是否随频率与可预测性发展出独特的表征。所有模型均显示出由频率与可预测性驱动的整体性存储证据,进一步支持基于使用的语言理论。
原文摘要 · Abstract (English)
One of the most central aspects of language processing is the ability to trade off between stored representations and abstract knowledge: one must retrieve stored representations, but also generate novel ones by applying productive rules. While recent work has examined abstract knowledge in language models, holistic storage has received far less attention. We probe internal representations in both text-based LLMs and an ASR model, testing whether V+up phrasal verbs develop distinct representations as a function of frequency and predictability. All models show evidence of holistic storage driven by frequency and predictability, further supporting usage-based theories of language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。