arXiv:2604.10401cs.CL2026-04中稿 · the 39th Canadian …

用大模型扩充数据,让名字猜国籍更准更快。

NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data

论文配图:NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data
图 1 · 摘自论文原文
  • 用大模型生成低资源国家姓名,扩增训练数据
  • 在合成尾部数据上准确率提升显著,真实数据也小幅提高
  • 适合需要大规模、高精度国籍识别的科研与合规场景

从个人姓名推断国籍是公平性监控、个性化推荐及生物医学与社会学研究中的关键能力。现有基于姓名的国籍分类器多依赖小规模或特定来源的标注数据,导致覆盖不足且对少数国家表现差。虽然大语言模型(LLMs)在零样本下表现优异,但其计算成本和延迟使其难以用于实时大规模部署。本文利用开放学术图谱(OAG)构建大规模姓名-国籍数据集,并提出一种框架:将LLM作为数据增强工具而非推理引擎。通过生成低资源国家的姓名来扩充数据,在真实与合成尾部测试集上评估。结果显示,当评估包含合成尾部姓名时,数据增强带来显著性能提升;即使在普通尾部国家指标上也有小幅改善。整体上,NameBERT模型在跨域与非跨域任务中均显著优于现有最佳基线,且比直接使用LLM更高效,适合大规模推理。

原文摘要 · Abstract (English)

Inferring nationality from personal names is a critical capability for equity and bias monitoring, personalization, and a valuable tool in biomedical and sociological research. However, existing name-based nationality classifiers are typically trained on relatively small or source-specific labeled datasets, which can introduce coverage gaps and limit performance for underrepresented countries. While large language models (LLMs) demonstrate strong zero-shot performance for name-based nationality prediction, their computational cost and latency make them impractical for real-time, large-scale deployment. In this work, we created a large-scale name-nationality dataset from the Open Academic Graph (OAG) and introduce a framework that leverages LLMs as dataset enrichers rather than inference engines. We augment low-resource countries with LLM-generated names and evaluate on real and synthetic-tail test sets. We find that augmentation produces large gains when evaluation includes synthetic tail names and still offers a modest lift on tail-country metrics otherwise. Overall, NameBERT models achieve significantly higher accuracy than state-of-the-art baselines across both in- and out-of-domain tasks, while remaining efficient for large-scale inference compared to LLMs.

国籍识别数据增强大模型应用命名实体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。