对比神经模型与大模型预测姓名国籍,发现大模型更准且懂常识。
Nationality and Region Prediction from Names: A Comparative Study of Neural Models and Large Language Models
- 用大模型和神经网络比对姓名国籍预测,分三级粒度评估。
- 大模型在所有级别都更优,尤其对低频国籍表现更好。
- 适合做跨文化分析、需理解地理关联的研究者参考。
从姓名预测国籍在营销、人口研究和家谱学中具有实际价值。传统神经模型依赖特定数据学习姓名与国籍的统计关联,难以泛化至低频国籍或区分同区域相似国籍。大语言模型(LLMs)可通过预训练获得的世界知识缓解这些问题。本研究系统比较了六种神经模型与六种LLM提示策略,在国籍、区域、大陆三个粒度层级上进行评估,结合频率分层分析与错误分析。结果表明,LLMs在所有粒度层级均优于神经模型,且粒度越粗差距越小。简单机器学习方法对低频类别最鲁棒,而预训练模型与LLMs在低频情况下性能下降。错误分析显示,LLMs常出现“近似错误”——国籍错但区域正确;神经模型则更多出现跨区域误判,且偏向高频类别。研究揭示,LLM优势源于世界知识,模型选择应考虑所需粒度,评估也应关注错误质量而非仅准确率。
原文摘要 · Abstract (English)
Predicting nationality from personal names has practical value in marketing, demographic research, and genealogical studies. Conventional neural models learn statistical correspondences between names and nationalities from task-specific training data, posing challenges in generalizing to low-frequency nationalities and distinguishing similar nationalities within the same region. Large language models (LLMs) have the potential to address these challenges by leveraging world knowledge acquired during pre-training. In this study, we comprehensively compare neural models and LLMs on nationality prediction, evaluating six neural models and six LLM prompting strategies across three granularity levels (nationality, region, and continent), with frequency-based stratified analysis and error analysis. Results show that LLMs outperform neural models at all granularity levels, with the gap narrowing as granularity becomes coarser. Simple machine learning methods exhibit the highest frequency robustness, while pre-trained models and LLMs show degradation for low-frequency nationalities. Error analysis reveals that LLMs tend to make ``near-miss'' errors, predicting the correct region even when nationality is incorrect, whereas neural models exhibit more cross-regional errors and bias toward high-frequency classes. These findings indicate that LLM superiority stems from world knowledge, model selection should consider required granularity, and evaluation should account for error quality beyond accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。