arXiv:2409.11491cs.CL2024-09被引 6

用大模型零样本预测姓名背后的性别种族等信息,效果优于传统方法。

Enriching Datasets with Demographics through Large Language Models: What's in a Name?

  • 直接利用大模型零样本能力预测姓名对应的种族、性别等信息。
  • 在多个数据集上表现优于专门训练的模型,包括香港金融从业者真实数据。
  • 揭示了大模型在人口统计预测中的固有偏见,推动后续去偏研究。

从姓名中推断性别、种族、年龄等人口统计信息,是医疗、公共政策与社会科学中的关键任务。尽管已有研究采用隐马尔可夫模型和循环神经网络进行预测,但仍存在缺乏大规模、高质量、无偏见且公开可用的数据集,以及跨数据集鲁棒性差的问题,制约了传统监督学习的发展。本文证明,大型语言模型(LLMs)的零样本能力在该任务上可达到甚至超过专门训练模型的表现。我们将其应用于多个数据集,包括一份香港持牌金融专业人士的真实未标注数据,并深入评估模型内在的人口统计偏见。本工作不仅推进了人口统计信息扩充的技术前沿,也为未来缓解大模型偏见提供了新方向。

原文摘要 · Abstract (English)

Enriching datasets with demographic information, such as gender, race, and age from names, is a critical task in fields like healthcare, public policy, and social sciences. Such demographic insights allow for more precise and effective engagement with target populations. Despite previous efforts employing hidden Markov models and recurrent neural networks to predict demographics from names, significant limitations persist: the lack of large-scale, well-curated, unbiased, publicly available datasets, and the lack of an approach robust across datasets. This scarcity has hindered the development of traditional supervised learning approaches. In this paper, we demonstrate that the zero-shot capabilities of Large Language Models (LLMs) can perform as well as, if not better than, bespoke models trained on specialized data. We apply these LLMs to a variety of datasets, including a real-life, unlabelled dataset of licensed financial professionals in Hong Kong, and critically assess the inherent demographic biases in these models. Our work not only advances the state-of-the-art in demographic enrichment but also opens avenues for future research in mitigating biases in LLMs.

大模型应用数据增强偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。