用19世纪美国文学方言数据集检验语言模型对拼写变异的感知能力。
Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus
- 构建带方言标签的文学拼写变异数据集,支持计算实验。
- 不同分词方式显著影响模型识别拼写信息的能力。
- 揭示拼写变异通过多语言通道传递语义,适合方言研究者。
我们提出一个包含19世纪美国文学正字变异词元的数据集,并添加了人类标注的方言群体标签,旨在为探索具有文学意义的正字变异提供计算实验基础。基于该数据集,我们使用词级(BERT)和字符级(CANINE)上下文语言模型进行了初步广泛实验。结果表明,由有意拼写变异产生的“方言效应”涉及多种语言通道,且这些通道在不同语言建模假设下可被不同程度地揭示。具体而言,分词方案的选择显著影响模型所能提取的拼写信息类型。
原文摘要 · Abstract (English)
We present a dataset of 19th century American literary orthovariant tokens with a novel layer of human-annotated dialect group tags designed to serve as the basis for computational experiments exploring literarily meaningful orthographic variation. We perform an initial broad set of experiments over this dataset using both token (BERT) and character (CANINE)-level contextual language models. We find indications that the "dialect effect" produced by intentional orthographic variation employs multiple linguistic channels, and that these channels are able to be surfaced to varied degrees given particular language modelling assumptions. Specifically, we find evidence showing that choice of tokenization scheme meaningfully impact the type of orthographic information a model is able to surface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。