arXiv:2409.18724cs.IRcs.CL2024-09

通过共现模式实现跨领域关键词提取,提升准确率与泛化能力。

Cross-Domain Keyword Extraction with Keyness Patterns

  • 基于独立与依赖特征构建关键词重要性模式
  • 跨领域测试中平均前10名F1达0.346,优于现有方法
  • 适合需要低标注依赖的跨领域文本分析场景

领域依赖性和标注主观性给监督式关键词提取带来挑战。本文基于二阶关键词重要性模式在社区层面存在且可从标注数据集中学习的前提,提出一种监督排序方法。该方法利用独立特征(如子领域、词长)和三类依赖特征(启发式、特异性、代表性)构成关键词重要性模式。采用两个基于卷积神经网络的模型从关键词数据集中学习这些模式,并通过自举采样策略克服标注主观性。实验表明,该方法在十项标准关键词数据集上达到当前最优性能,平均前10名F1为0.316;在四个未参与训练的数据集上仍保持0.346的平均前10名F1,展现出强跨领域鲁棒性。这种鲁棒性源于社区级关键模式数量有限、对语言领域不敏感,以及独立与依赖特征的区分机制,结合采样训练策略有效平衡了过拟合风险与负样本不足问题。

原文摘要 · Abstract (English)

Domain dependence and annotation subjectivity pose challenges for supervised keyword extraction. Based on the premises that second-order keyness patterns are existent at the community level and learnable from annotated keyword extraction datasets, this paper proposes a supervised ranking approach to keyword extraction that ranks keywords with keyness patterns consisting of independent features (such as sublanguage domain and term length) and three categories of dependent features -- heuristic features, specificity features, and representavity features. The approach uses two convolutional-neural-network based models to learn keyness patterns from keyword datasets and overcomes annotation subjectivity by training the two models with bootstrap sampling strategy. Experiments demonstrate that the approach not only achieves state-of-the-art performance on ten keyword datasets in general supervised keyword extraction with an average top-10-F-measure of 0.316 , but also robust cross-domain performance with an average top-10-F-measure of 0.346 on four datasets that are excluded in the training process. Such cross-domain robustness is attributed to the fact that community-level keyness patterns are limited in number and temperately independent of language domains, the distinction between independent features and dependent features, and the sampling training strategy that balances excess risk and lack of negative training data.

关键词提取跨领域深度学习模式识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。