arXiv:2503.00958cs.CL2025-03EMNLP被引 3

利用Transformer各层特征分析作者风格,提升跨领域识别效果。

Layered Insights: Generalizable Analysis of Authorial Style by Leveraging All Transformer Layers

  • 融合预训练模型多层语言表征进行作者识别
  • 跨域测试下准确率超越现有最佳方法
  • 揭示不同层对特定风格特征的专门化作用

我们提出一种新的作者归属方法,利用预训练Transformer模型在不同层次学习到的语言表征。在三个数据集上评估该方法,对比了域内与域外场景下的最先进基线模型。结果表明,结合多个Transformer层能显著提升作者归属模型在域外数据上的鲁棒性,取得新的最先进性能。进一步分析显示,模型不同层在表征特定风格特征方面具有专业化倾向,这有助于模型在域外测试时表现更优。

原文摘要 · Abstract (English)

We propose a new approach for the authorship attribution task that leverages the various linguistic representations learned at different layers of pre-trained transformer-based models. We evaluate our approach on three datasets, comparing it to a state-of-the-art baseline in in-domain and out-of-domain scenarios. We found that utilizing various transformer layers improves the robustness of authorship attribution models when tested on out-of-domain data, resulting in new state-of-the-art results. Our analysis gives further insights into how our model's different layers get specialized in representing certain stylistic features that benefit the model when tested out of the domain.

作者识别Transformer风格分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。