arXiv:2409.13706cs.CL2024-09被引 1

用拼音或粤拼表示中文姓名可提升数据链接准确率。

Decolonising Data Systems: Using Jyutping or Pinyin as tonal representations of Chinese names for data linkage

  • 采用带声调的拼音或粤拼替代非标准罗马化
  • 实验显示拼音和粤拼错误率低于港府罗马化系统
  • 适合关注数据公平性与跨文化研究的学者

数据链接在健康研究与政策制定中日益重要,但其有效性依赖于数据质量,而不同群体的链接率差异可能引入选择偏差。姓名罗马化是影响数据质量的关键因素。将汉字等字符系统的名称转换为拉丁字母时,若未标准化且缺乏声调信息,会显著降低华裔移民的链接成功率。本文建议采用包含声调信息的标准罗马化方案——粤拼(Jyutping)或普通话拼音(Pinyin),以改善中文姓名的数据链接效果。通过分析771个公开获取的中英文姓名,对比了粤拼、拼音与香港政府罗马化系统(HKG-romanisation)的表现,结果表明:粤拼与拼音均比港府系统产生更少错误。文章强调应以原始书写形式收集并保存个人姓名,这在伦理和社会层面具有重要意义。该做法可推动针对语言特性的预处理与链接范式发展,使研究数据更具包容性,更好反映目标人群。

原文摘要 · Abstract (English)

Data linkage is increasingly used in health research and policy making and is relied on for understanding health inequalities. However, linked data is only as useful as the underlying data quality, and differential linkage rates may induce selection bias in the linked data. A mechanism that selectively compromises data quality is name romanisation. Converting text of a different writing system into Latin based writing, or romanisation, has long been the standard process of representing names in character-based writing systems such as Chinese, Vietnamese, and other languages such as Swahili. Unstandardised romanisation of Chinese characters, due in part to problems of preserving the correct name orders the lack of proper phonetic representation of a tonal language, has resulted in poor linkage rates for Chinese immigrants. This opinion piece aims to suggests that the use of standardised romanisation systems for Cantonese (Jyutping) or Mandarin (Pinyin) Chinese, which incorporate tonal information, could improve linkage rates and accuracy for individuals with Chinese names. We used 771 Chinese and English names scraped from openly available sources, and compared the utility of Jyutping, Pinyin and the Hong Kong Government Romanisation system (HKG-romanisation) for representing Chinese names. We demonstrate that both Jyutping and Pinyin result in fewer errors compared with the HKG-romanisation system. We suggest that collecting and preserving people's names in their original writing systems is ethically and socially pertinent. This may inform development of language-specific pre-processing and linkage paradigms that result in more inclusive research data which better represents the targeted populations.

数据链接中文姓名声调信息公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。