首个标注了波斯英语混用词词性的大规模语料库,助力多语言社交文本分析。
PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

- 构建混合语言标注框架,利用大模型辅助生成词性标签与主题
- 6800条社交平台数据中,名词是混用词主要类别,平台间分布差异显著
- 适用于多语言自然语言处理、社会语言学研究者
社交媒体成为多语言交流的重要场所,用户常在单条语句中混合使用多种语言。尽管已有多个语言对的混用语料库,但波斯语-英语混用仍相对未被充分研究。现有波斯语资源缺乏通用依存关系(UD)标注的混用词词性信息,限制了语言学分析与语法感知型NLP模型的发展。为此,我们推出PERCEPT,首个公开可用的大规模波斯语-英语混用语料库,包含代码混用词的通用依存关系词性标注。该数据集由来自X、Instagram和Digikala的6,800条帖子构成。我们还提出一种大模型辅助的标注框架,可自动分配词性标签与文档级主题。人工评估显示,自动生成标注与人工标注高度一致,验证了标注可靠性。基于PERCEPT,我们首次在多个社交平台对波斯语-英语混用进行综合语言学分析。结果表明,名词是混用词的主要类别,其他词性分布随平台而异。混用词的位置分布跨平台高度一致,但在Digikala中触发效应尤为明显。PERCEPT已公开于https://github.com/kalhorghazal/PERCEPT。
原文摘要 · Abstract (English)
Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs, Persian-English code-mixing remains relatively underexplored. Existing Persian resources lack Universal Dependencies (UD) part-of-speech (POS) annotations for code-mixed words, limiting both linguistic analyses and the development of syntax-aware NLP models. To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words. The dataset comprises 6,800 posts collected from X, Instagram, and Digikala. We further present an LLM-assisted annotation framework that automatically assigns POS tags and document-level topics. Human evaluation demonstrates high agreement between the automatically generated annotations and gold annotations, confirming the reliability of the annotations. Using PERCEPT, we conduct the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms. Our analyses reveal that nouns are the predominant category for code-mixed words, while the distributions of other POS categories vary across platforms. We further find that the positional distribution of code-mixed words is remarkably consistent across platforms, whereas the triggering effect is substantially more pronounced in Digikala. PERCEPT is publicly available at https://github.com/kalhorghazal/PERCEPT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。