针对低资源语言数据稀缺,提出改善数据收集与标注的伦理与质量建议。
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce

- 通过调研直接相关者,发现数据质量与文化适切性问题。
- 揭示参与式研究被误用等标注实践中的伦理风险。
- 为尊重语者尊严与劳动,提出更人本的数据开发指南。
语言是一种象征资本,深刻影响人们的生活(Bourdieu1977,1991)。作为沟通工具,它反映身份、文化、传统与社会结构。因此,一种语言的数据不应仅被视为词元集合。严谨的数据收集与标注对发展更具人文关怀和社会意识的技术至关重要。尽管自然语言处理领域对低资源语言的兴趣日益增长,但该领域仍面临数据稀缺与合格标注者匮乏等独特挑战。本文收集了直接参与或受NLP成果影响的中低资源语言使用者反馈,结合定量与定性分析,揭示两大核心问题:(1) 数据质量,包括语言与文化的适切性;(2) 常见标注实践中的伦理问题,如参与式研究的滥用。基于此,我们提出多项建议,旨在创建反映语者文化背景的高质量语言资源,同时尊重数据工作者的尊严与劳动。
原文摘要 · Abstract (English)
Language is a form of symbolic capital that affects people's lives in many ways (Bourdieu1977,1991). As a powerful means of communication, it reflects identities, cultures, traditions, and societies more broadly. Therefore, data in a given language should be regarded as more than just a collection of tokens. Rigorous data collection and labeling practices are essential for developing more human-centered and socially aware technologies. Although there has been growing interest in under-resourced languages within the NLP community, work in this area faces unique challenges, such as data scarcity and limited access to qualified annotators. In this paper, we collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages. We conduct both quantitative and qualitative analyses of their responses and highlight key issues related to: (1) data quality, including linguistic and cultural appropriateness; and (2) the ethics of common annotation practices, such as the misuse of participatory research. Based on these findings, we make several recommendations for creating high-quality language artefacts that reflect the cultural milieu of their speakers, while also respecting the dignity and labor of data workers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。