研究发现,句子越可预测,越可能省略'that',且新方法更准确捕捉这种规律。
Uniform Information Density and Syntactic Reduction: Revisiting $\textit{that}$-Mentioning in English Complement Clauses
- 用上下文词嵌入估算信息密度,比旧法更准
- 可预测性高的从句中,'that'省略概率更高
- 适合对语言生成机制和认知模型感兴趣者
说话者常有多种表达相同意义的方式。统一信息密度(UID)假说认为,说话人会利用这种变异性,在语言产出中保持信息传递速率的稳定。基于先前将UID与句法简化关联的研究,我们重新考察了英语补语从句中可选助词'that'的省略现象:当从句信息密度低(即更可预测)时,'that'更可能被省略。本研究利用大规模当代对话语料库,结合机器学习与神经语言模型,改进了信息密度的估算方法。结果复现了信息密度与'that'使用之间的既有关系。但发现,以往基于主句动词子范畴概率的信息密度测量方式,存在显著的词汇特异性偏差;而基于上下文词嵌入的估计则能解释更多关于助词使用模式的变异。
原文摘要 · Abstract (English)
Speakers often have multiple ways to express the same meaning. The Uniform Information Density (UID) hypothesis suggests that speakers exploit this variability to maintain a consistent rate of information transmission during language production. Building on prior work linking UID to syntactic reduction, we revisit the finding that the optional complementizer $\textit{that}$ in English complement clauses is more likely to be omitted when the clause has low information density (i.e., more predictable). We advance this line of research by analyzing a large-scale, contemporary conversational corpus and using machine learning and neural language models to refine estimates of information density. Our results replicated the established relationship between information density and $\textit{that}$-mentioning. However, we found that previous measures of information density based on matrix verbs' subcategorization probability capture substantial idiosyncratic lexical variation. By contrast, estimates derived from contextual word embeddings account for additional variance in patterns of complementizer usage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。