arXiv:2609.06771cs.CLcs.AI2026-09

构建首个跨语言跨文体大规模作者风格识别基准,助力文本身份分析研究。

AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths

论文配图:AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths
图 1 · 摘自论文原文
  • 设计多语言、多文体、多长度的统一评估框架,覆盖10语言9主类66细类
  • 零样本测试下最优模型检索准确率仅25.8%,验证任务最佳AUC达96.8%
  • 揭示不同模型在任务间表现差异,适合研究作者风格泛化与失效机制

作者特征在数字取证、抄袭检测、账号关联、虚假信息调查及机器生成文本识别等场景中至关重要。然而现有作者识别基准分散,通常仅覆盖单一语言、一种文体或有限文档长度,难以评估现代表示方法的泛化能力。本文提出AuthBench,一个大规模多语言作者表示基准,涵盖10种主流语言、9大主要文体、66个细粒度文体类别及4种文档长度区间,共包含428,150篇由153,825名作者撰写的文档。支持两类互补任务:作者归属(同作者检索)和作者验证(同作者二分类判断)。我们在统一零样本协议下对47个神经网络模型和3个非神经基线进行评测。结果显示,作者表示仍远未解决:最优检索模型成功@5仅为0.258,最优验证模型达到0.076 EER和0.968 ROC-AUC。排行榜还揭示任务间的显著差异,不同模型族在检索与验证任务上表现各异,且在语言、文体和长度维度上存在明显性能差距。这些发现使AuthBench不仅是新基准,更成为诊断作者表示成功或失败条件的工具。数据与工具已开源至https://github.com/mao-code/AuthBench 和 https://huggingface.co/datasets/MaoXun/AuthBench。

原文摘要 · Abstract (English)

Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism analysis, account linking, misinformation investigation, and machine-generated text detection. Yet current authorship benchmarks remain fragmented, usually covering only a narrow language set, a single genre, or a limited document-length regime, which makes it difficult to assess whether modern representations truly generalize. We introduce AuthBench, a large-scale multilingual benchmark for authorship representation that is designed to make this evaluation broad, standardized, and realistic. AuthBench contains 428,150 documents written by 153,825 individuals across ten widely used languages, 9 primary genres, 66 fine-grained genres, and four document-length buckets. It supports two complementary tasks: authorship attribution, formulated as same-author retrieval and authorship verification, formulated as same-author binary decision. We benchmark 47 neural models and three non-neural baselines under a unified zero-shot protocol. Results show that authorship representation remains far from solved: the best retrieval model reaches only 0.258 Success@5, while the best verification model achieves 0.076 EER and 0.968 ROC-AUC. The leaderboard also reveals a meaningful task split, with different model families leading retrieval and verification, and large performance differences across languages, genres, and lengths. These findings position AuthBench not only as a new benchmark, but as a diagnostic resource for studying when and why authorship representations succeed or fail. We release AuthBench, its evaluation toolkit, and benchmark data at https://github.com/mao-code/AuthBench and https://huggingface.co/datasets/MaoXun/AuthBench.

作者识别多语言基准测试文本分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。