arXiv:2504.08527cs.CL2025-04被引 1

融合BERT与传统特征的集成模型显著提升日文文学作品作者归属准确率

Integrated ensemble of BERT- and features-based models for authorship attribution in Japanese literary works

  • 构建BERT与传统特征的集成模型,结合两者优势
  • 在小样本数据下,集成模型F1分数提升约14点
  • 适合需要高精度作者归属的小规模文本分析任务

传统作者归属(AA)任务依赖文本风格特征的统计分析与分类。近年来,预训练语言模型(PLMs)在短文本分类中表现优异,但在小样本场景下,尤其在AA任务中的效果仍不明确。此外,如何有效结合PLM与传统特征方法仍具挑战。本研究旨在通过融合传统特征与现代PLM方法的集成模型,显著提升小样本下的日文文学作品作者归属性能。实验使用两个文学作品语料库,分别对10位作者进行分类。结果表明,即使在小样本下,BERT仍具有效性;基于BERT的模型及分类器集成均优于单一模型,而集成方法进一步显著提升性能。对于未包含于预训练数据的语料库,集成模型相比最优单模型,F1分数提升约14点。该方法为未来高效利用日益增长的数据处理工具提供了可行方案。

原文摘要 · Abstract (English)

Traditionally, authorship attribution (AA) tasks relied on statistical data analysis and classification based on stylistic features extracted from texts. In recent years, pre-trained language models (PLMs) have attracted significant attention in text classification tasks. However, although they demonstrate excellent performance on large-scale short-text datasets, their effectiveness remains under-explored for small samples, particularly in AA tasks. Additionally, a key challenge is how to effectively leverage PLMs in conjunction with traditional feature-based methods to advance AA research. In this study, we aimed to significantly improve performance using an integrated integrative ensemble of traditional feature-based and modern PLM-based methods on an AA task in a small sample. For the experiment, we used two corpora of literary works to classify 10 authors each. The results indicate that BERT is effective, even for small-sample AA tasks. Both BERT-based and classifier ensembles outperformed their respective stand-alone models, and the integrated ensemble approach further improved the scores significantly. For the corpus that was not included in the pre-training data, the integrated ensemble improved the F1 score by approximately 14 points, compared to the best-performing single model. Our methodology provides a viable solution for the efficient use of the ever-expanding array of data processing tools in the foreseeable future.

作者归属BERT小样本集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。