从对话数据中自动提取语言单元,揭示新兴语言的结构规律。
Morpheme Induction for Emergent Language
- 基于互信息贪心筛选形式与语义匹配的词素对
- 在生成数据和人类语言上验证,能准确预测词素
- 适用于分析新兴语言的同义、多义等特征
我们提出CSAR算法,用于从并行话语与语义数据中诱导词素。该算法为贪心策略,依次执行:(1)根据形式与意义的互信息计算词素权重;(2)选取权重最高的词素对;(3)将其从语料中移除;(4)重复过程以诱导更多词素(即:计数、选择、剔除、重复)。首先在程序生成的数据集上验证了CSAR的有效性,并与相关任务基线对比。其次,在人类语言数据上验证其性能,表明算法在相邻领域具有合理预测能力。最后,对若干新兴语言进行分析,量化其同义度与多义度等语言特征。
原文摘要 · Abstract (English)
We introduce CSAR, an algorithm for inducing morphemes from emergent language corpora of parallel utterances and meanings. It is a greedy algorithm that (1) weights morphemes based on mutual information between forms and meanings, (2) selects the highest-weighted pair, (3) removes it from the corpus, and (4) repeats the process to induce further morphemes (i.e., Count, Select, Ablate, Repeat). The effectiveness of CSAR is first validated on procedurally generated datasets and compared against baselines for related tasks. Second, we validate CSAR's performance on human language data to show that the algorithm makes reasonable predictions in adjacent domains. Finally, we analyze a handful of emergent languages, quantifying linguistic characteristics like degree of synonymy and polysemy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。