用子词嵌入检测卢森堡语评论中的拼写与形态变异,无需预处理或人工标注。
A Subword Embedding Approach for Variation Detection in Luxembourgish User Comments
- 在原始文本上训练子词嵌入,结合余弦与n-gram相似度聚类相关形式。
- 发现大量词汇与正字法变异,符合方言和社会语言学研究中的模式。
- 方法无需人工标注,结果可解释,适合小语种和多语言场景研究。
本文提出一种基于嵌入的变异检测方法,无需依赖预处理或预定义变体列表。该方法在原始文本上训练子词嵌入,并通过余弦相似度与n-gram相似度联合聚类相关形式,将拼写与形态多样性视为语言结构而非噪声进行分析。利用大规模卢森堡语用户评论语料库,该方法揭示了广泛存在的词汇与正字法变异,其模式与方言及社会语言学研究一致。生成的变体簇捕捉系统性对应关系,凸显区域与风格差异。该流程不严格依赖人工标注,但产生的聚类透明,支持定量与定性分析。结果表明,分布建模可在‘嘈杂’或低资源环境下揭示有意义的变异模式,为多语言及小语种中的语言多样性研究提供可复现的方法框架。
原文摘要 · Abstract (English)
This paper presents an embedding-based approach to detecting variation without relying on prior normalisation or predefined variant lists. The method trains subword embeddings on raw text and groups related forms through combined cosine and n-gram similarity. This allows spelling and morphological diversity to be examined and analysed as linguistic structure rather than treated as noise. Using a large corpus of Luxembourgish user comments, the approach uncovers extensive lexical and orthographic variation that aligns with patterns described in dialectal and sociolinguistic research. The induced families capture systematic correspondences and highlight areas of regional and stylistic differentiation. The procedure does not strictly require manual annotation, but does produce transparent clusters that support both quantitative and qualitative analysis. The results demonstrate that distributional modelling can reveal meaningful patterns of variation even in ''noisy'' or low-resource settings, offering a reproducible methodological framework for studying language variety in multilingual and small-language contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。