用机器学习分析英式口语中的年龄语言特征,可准确预测说话人年龄段。
Aged to Perfection: Machine-Learning Maps of Age in Conversational English
- 基于英国国家语料库2014,融合计算语言学与机器学习建模。
- 通过语句时长、词汇多样性等指标实现跨代际语言差异识别。
- 适用于研究社会语言学演化,或需年龄推断的语音分析场景。
本研究利用大型当代英式口语语料库——英国国家语料库2014(British National Corpus 2014),探讨不同年龄群体的语言模式差异。研究聚焦于说话人人口统计特征与语言因素(如语句持续时间、词汇多样性、词汇选择)之间的关联。通过结合计算语言分析与机器学习方法,我们旨在识别多代际特有的语言标记,并构建可从多个维度一致估计说话人年龄组的预测模型。该工作深化了对现代英式口语中社会语言多样性生命周期的理解。
原文摘要 · Abstract (English)
The study uses the British National Corpus 2014, a large sample of contemporary spoken British English, to investigate language patterns across different age groups. Our research attempts to explore how language patterns vary between different age groups, exploring the connection between speaker demographics and linguistic factors such as utterance duration, lexical diversity, and word choice. By merging computational language analysis and machine learning methodologies, we attempt to uncover distinctive linguistic markers characteristic of multiple generations and create prediction models that can consistently estimate the speaker's age group from various aspects. This work contributes to our knowledge of sociolinguistic diversity throughout the life of modern British speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。