提出首个德语作者身份验证大规模基准,提升跨域验证能力
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark
- 构建包含40万+文本对的德语作者验证基准GerAV,涵盖推特、红迪网多源数据
- 微调大模型在基准上实现最高0.09的F1提升,零样本超越GPT-5
- 揭示专精与泛化权衡,融合训练源可缓解跨场景性能下降
作者身份验证(AV)旨在判断两段文本是否出自同一作者,该任务在英语数据上研究充分,但其他语言的大规模基准和系统评估仍匮乏。本文提出GerAV,一个覆盖超过40万标注文本对的德语作者身份验证综合基准,数据源自推特与红迪网,其中红迪网部分进一步划分为同领域、跨领域消息类及资料档案类子集,支持对数据来源、主题领域和文本长度影响的受控分析。利用提供的训练划分,我们系统评估了强基线与前沿模型,发现最优方法——微调大语言模型——相比近期基线最高提升0.09绝对F1得分,并在零样本设置下比GPT-5高出0.08。进一步观察到:特定数据类型训练的模型在匹配条件下表现最佳,但跨数据环境泛化能力弱,该局限可通过融合不同训练源缓解。总体而言,GerAV为德语及跨域作者验证研究提供了一个挑战性强且多样化的基准。代码与数据访问信息已开源。
原文摘要 · Abstract (English)
Authorship verification (AV) is the task of determining whether two texts were written by the same author and has been studied extensively, predominantly for English data. In contrast, large-scale benchmarks and systematic evaluations for other languages remain scarce. We address this gap by introducing GerAV, a comprehensive benchmark for German AV comprising over 400k labeled text pairs. GerAV is built from Twitter and Reddit data, with the Reddit part further divided into in-domain and cross-domain message-based subsets, as well as a profile-based subset. This design enables controlled analysis of the effects of data source, topical domain, and text length. Using the provided training splits, we conduct a systematic evaluation of strong baselines and state-of-the-art models and find that our best approach, a fine-tuned large language model, outperforms recent baselines by up to 0.09 absolute F1 score and surpasses GPT-5 in a zero-shot setting by 0.08. We further observe a trade-off between specialization and generalization: models trained on specific data types perform best under matching conditions but generalize less well across data regimes, a limitation that can be mitigated by combining training sources. Overall, GerAV provides a challenging and versatile benchmark for advancing research on German and cross-domain AV. Our code and information about data access are available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。