发现并抑制文本识别噪声敏感神经元,提升古籍实体识别准确率
Investigating OCR-Sensitive Neurons to Improve Entity Recognition in Historical Documents
- 分析变压器模型中对文字噪声敏感的神经元激活模式
- 在古籍报纸与注疏数据集上实现命名实体识别性能提升
- 适合关注历史文本处理与模型鲁棒性优化的研究者
本文研究了Transformer架构中是否存在对OCR噪声敏感的神经元,及其对历史文档命名实体识别(NER)性能的影响。通过分析模型在清晰与含噪文本输入下的神经元激活模式,识别并抑制了这些敏感神经元。基于两个开源大语言模型(Llama2和Mistral),实验验证了OCR敏感区域的存在,并在历史报纸与古典注疏数据集上显著提升了NER表现,表明针对特定神经元进行调控可有效增强模型在噪声文本上的性能。
原文摘要 · Abstract (English)
This paper investigates the presence of OCR-sensitive neurons within the Transformer architecture and their influence on named entity recognition (NER) performance on historical documents. By analysing neuron activation patterns in response to clean and noisy text inputs, we identify and then neutralise OCR-sensitive neurons to improve model performance. Based on two open access large language models (Llama2 and Mistral), experiments demonstrate the existence of OCR-sensitive regions and show improvements in NER performance on historical newspapers and classical commentaries, highlighting the potential of targeted neuron modulation to improve models' performance on noisy text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。