用AI从众包文献中自动提取关键词,发现无监督模型最适配且需重视数据责任。
Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI
- 对比命名实体、关键词抽取与主题建模三类NLP方法在众包数据中的表现。
- 传统无监督模型在准确率和可解释性上优于生成式AI,适合大规模应用。
- 强调数据治理责任,建议优先采用开源可解释模型以保障透明性。
为解决众包数字资源中大规模关键词标注的技术与伦理挑战,本文以牛津大学托管的“他们最辉煌的时刻”二战数字档案为案例,评估了命名实体识别、关键词提取与主题建模三类自然语言处理方法。研究涵盖从传统统计方法到现代生成式AI神经网络的多种技术路径。定量与定性分析表明,尽管各类NLP方法具备规模化提取潜力,但无单一方案能完全胜任;模型选择显著影响结果质量。文章指出,在由真实贡献者参与形成的众包资源中,自动化关键词提取不仅关乎技术性能,更涉及数据治理责任。评估结果显示,开源可解释的抽取式模型更适合负责任部署,而生成式AI虽具抽象能力,却带来问责风险,需审慎权衡。
原文摘要 · Abstract (English)
Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections. This article reports the findings of the "Extracting Keywords from Crowdsourced Collections" project, which used the Their Finest Hour Online Archive, a crowdsourced Second World War digital collection hosted by the University of Oxford, as a case study. The project evaluated three Natural Language Processing approaches to automate keyword extraction: Named Entity Recognition, Keyword Extraction, and Topic Modelling. It tested these approaches across a range of artificial intelligence techniques, from traditional statistical methods to modern GenAI neural networks. Our quantitative and qualitative findings indicate that Natural Language Processing approaches offer real potential for keyword extraction at scale in crowdsourced collections, but that no single method offers a complete solution and that model choice significantly shapes results. We argue that in crowdsourced collections, where metadata is the direct product of engagement with living contributors, automated keyword extraction raises distinct stewardship responsibilities that must be addressed alongside technical performance. Open-weight, extractive models emerge from our evaluation as best placed to support responsible deployment, while generative AI, despite its abstractive potential, introduces accountability risks that anyone managing crowdsourced collections should weigh carefully.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。