arXiv:2505.24538cs.CL2025-05ACL被引 1

用AI识别文化遗产数据中的冒犯性术语并提供背景说明,不删除只揭示。

Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage Collections

  • 结合多方共建的多语言词汇库与大模型,检测文化遗产元数据中的敏感词。
  • 已处理超790万条记录,为争议性术语提供历史与当下认知背景。
  • 适合文博机构、数字人文研究者,推动包容性文化数据建设。

文化遗产数据蕴含社会历史、传统与身份的重要知识,塑造我们对过去与现在的理解。然而,许多文化遗产藏品包含反映历史偏见的过时或冒犯性描述。文化机构面临巨大挑战,因数据规模庞大且复杂。为此,我们开发了一款人工智能工具,可检测文化遗产元数据中的冒犯性术语,并提供其历史背景与当代认知的上下文信息。该工具融合了由边缘群体、研究人员与文博专业人员共同构建的多语言词汇库,结合传统自然语言处理技术与大型语言模型(LLMs)。工具以独立网页应用形式发布,并集成至多个主流文化遗产平台,已处理超过790万条记录,对其中检测到的争议性术语进行语境化阐释。本方法不主张删除这些术语,而是通过揭示偏见,提供可操作洞察,助力构建更具包容性与可访问性的文化遗产收藏。

原文摘要 · Abstract (English)

Cultural Heritage (CH) data hold invaluable knowledge, reflecting the history, traditions, and identities of societies, and shaping our understanding of the past and present. However, many CH collections contain outdated or offensive descriptions that reflect historical biases. CH Institutions (CHIs) face significant challenges in curating these data due to the vast scale and complexity of the task. To address this, we develop an AI-powered tool that detects offensive terms in CH metadata and provides contextual insights into their historical background and contemporary perception. We leverage a multilingual vocabulary co-created with marginalized communities, researchers, and CH professionals, along with traditional NLP techniques and Large Language Models (LLMs). Available as a standalone web app and integrated with major CH platforms, the tool has processed over 7.9 million records, contextualizing the contentious terms detected in their metadata. Rather than erasing these terms, our approach seeks to inform, making biases visible and providing actionable insights for creating more inclusive and accessible CH collections.

文化遗产AI伦理敏感词检测大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。