arXiv:2503.01217cs.CLcs.LG2025-03中稿 · publication at the…被引 5

改进中文命名实体识别边界划分与语义表示,提升准确率。

HREB-CRF: Hierarchical Reduced-bias EMA for Chinese Named Entity Recognition

  • 分层注意力加权平均缓解边界偏差问题。
  • 在多个数据集上F1值提升1.1%至9.8%。
  • 适合需要高精度实体识别的中文NLP任务。

中文命名实体识别常因边界划分错误、语义表征复杂及发音意义差异导致错误。本文提出HREB-CRF框架:基于分层减偏指数固定偏差加权平均的注意力机制结合条件随机场。该方法通过增强词边界并聚合长文本梯度,在MSRA、Resume和Weibo数据集上的实验显示F1值分别优于基线模型1.1%、1.6%和9.8%。显著的性能提升证明了该方法在中文命名实体识别任务中的有效性与鲁棒性。

原文摘要 · Abstract (English)

Incorrect boundary division, complex semantic representation, and differences in pronunciation and meaning often lead to errors in Chinese Named Entity Recognition(CNER). To address these issues, this paper proposes HREB-CRF framework: Hierarchical Reduced-bias EMA with CRF. The proposed method amplifies word boundaries and pools long text gradients through exponentially fixed-bias weighted average of local and global hierarchical attention. Experimental results on the MSRA, Resume, and Weibo datasets show excellent in F1, outperforming the baseline model by 1.1\%, 1.6\%, and 9.8\%. The significant improvement in F1 shows evidences of strong effectiveness and robustness of approach in CNER tasks.

命名实体识别中文NLP注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。