arXiv:2502.03292cs.CLcs.AI2025-02被引 3

用主动学习+模式训练,少标注数据也能高效识别维基需引用句

ALPET: Active Few-shot Learning for Citation Worthiness Detection in Low-Resource Wikipedia Languages

  • 结合主动学习与模式训练,自动筛选最有价值的标注样本
  • 仅需300个标注样本即达性能峰值,比基线减少超80%标注量
  • 适合资源匮乏语言的维基内容可信度提升,尤其适合小语种

引文必要性检测(CWD)旨在判断文章中哪些句子需要引用以验证信息。本文提出ALPET框架,融合主动学习(AL)与模式挖掘训练(PET),提升低资源语言下的CWD性能。在加泰罗尼亚语、巴斯克语和阿尔巴尼亚语维基数据集上,ALPET优于现有CCW基线,部分情况下标注量减少超过80%。性能在300个标注样本后趋于稳定,表明其适用于缺乏大规模标注数据的场景。尽管基于K-Means聚类的主动学习策略有一定优势,但其增益有限,尤其在小数据集上,常不及随机采样。这表明在资源受限环境下,随机采样仍是可靠基线。总体而言,ALPET以极少标注实现高精度,是提升低资源语言维基内容可验证性的有力工具。

原文摘要 · Abstract (English)

Citation Worthiness Detection (CWD) consists in determining which sentences, within an article or collection, should be backed up with a citation to validate the information it provides. This study, introduces ALPET, a framework combining Active Learning (AL) and Pattern-Exploiting Training (PET), to enhance CWD for languages with limited data resources. Applied to Catalan, Basque, and Albanian Wikipedia datasets, ALPET outperforms the existing CCW baseline while reducing the amount of labeled data in some cases above 80\%. ALPET's performance plateaus after 300 labeled samples, showing it suitability for low-resource scenarios where large, labeled datasets are not common. While specific active learning query strategies, like those employing K-Means clustering, can offer advantages, their effectiveness is not universal and often yields marginal gains over random sampling, particularly with smaller datasets. This suggests that random sampling, despite its simplicity, remains a strong baseline for CWD in constraint resource environments. Overall, ALPET's ability to achieve high performance with fewer labeled samples makes it a promising tool for enhancing the verifiability of online content in low-resource language settings.

引文检测主动学习小语种低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。