用大模型分析双语对话中的话题与社会语言差异,发现性别和语言主导性影响表达方式。
Modeling Topics and Sociolinguistic Variation in Code-Switched Discourse: Insights from Spanish-English and Spanish-Guaraní
- 借助大模型自动标注3691句双语语料的话题、文体和语用功能。
- 揭示迈阿密数据中性别、语言主导性与话语功能的系统关联。
- 为低资源双语研究提供可复现的自动化分析方法,适合语言学与NLP交叉研究者。
本研究提出一种大语言模型辅助的标注流程,对两种语言类型差异显著的双语语境(西班牙语-英语与西班牙语-瓜拉尼语)中的社会语言学与话题特征进行分析。基于大语言模型,对总计3,691个代码切换句子进行了话题、体裁及语用功能的自动标注,整合了迈阿密双语语料库中的人口统计元数据,并为西班牙语-瓜拉尼语数据集新增话题标注。结果揭示迈阿密数据中性别、语言主导性与话语功能存在系统性关联,巴拉圭文本则呈现出正式瓜拉尼语与非正式西班牙语之间的明显双言分层。这些发现以语料规模的量化证据复现并拓展了早期交互与社会语言学观察。研究表明,大语言模型能够可靠地恢复传统上仅通过人工标注才能获得的可解释社会语言模式,推动了跨语言与低资源双语研究的计算方法发展。
原文摘要 · Abstract (English)
This study presents an LLM-assisted annotation pipeline for the sociolinguistic and topical analysis of bilingual discourse in two typologically distinct contexts: Spanish-English and Spanish-Guaraní. Using large language models, we automatically labeled topic, genre, and discourse-pragmatic functions across a total of 3,691 code-switched sentences, integrated demographic metadata from the Miami Bilingual Corpus, and enriched the Spanish-Guaraní dataset with new topic annotations. The resulting distributions reveal systematic links between gender, language dominance, and discourse function in the Miami data, and a clear diglossic division between formal Guaraní and informal Spanish in Paraguayan texts. These findings replicate and extend earlier interactional and sociolinguistic observations with corpus-scale quantitative evidence. The study demonstrates that large language models can reliably recover interpretable sociolinguistic patterns traditionally accessible only through manual annotation, advancing computational methods for cross-linguistic and low-resource bilingual research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。