arXiv:2609.02379cs.CLcs.AI2026-09

构建多语言长文本伪造检测基准,评估模型在跨领域时的识别能力。

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

论文配图:MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts
图 1 · 摘自论文原文
  • 基于五款主流大模型生成928本跨六语种长文本,平均每本约5.9万字。
  • 现有方法在分布偏移下性能普遍下降,无一始终领先。
  • 适合研究鲁棒性伪造检测的学者和安全团队使用。

当前大语言模型作者归属(AA)研究的基准数据集仍受限,多数仅覆盖英语、特定场景或较旧模型,且多为短文本。本文提出MultiGhostBench,一个包含928本由五款近期大模型生成的多语言书籍的基准,涵盖六种语言与三种文字系统,平均每本书约5.9万字。该基准支持在领域、作者和语言分布偏移下的评估。对代表性AA方法的测试表明,无单一方法在所有场景下表现最优,且性能普遍随分布偏移下降。基于Transformer的检测器可在不同语言间保留生成器特征,但跨语言迁移效果因语言对而异;统计与指纹类方法则更依赖具体语言。我们期望MultiGhostBench能推动鲁棒性AA方法的发展与评估。数据与代码已开源:https://github.com/GrecoMT/MultiGhostBench。

原文摘要 · Abstract (English)

While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce MultiGhostBench, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book. The benchmark supports evaluation under domain, author, and language shifts. Evaluation of representative AA methods shows that no single method consistently performs best across settings, and performance generally degrades under distribution shifts. Transformer-based detectors can retain generator-related information across languages, although transfer effectiveness varies by language pair, whereas statistical and fingerprint-based detectors are more language-dependent. We envision MultiGhostBench as a valuable resource for the development and evaluation of robust AA methods. The dataset and code can be found at https://github.com/GrecoMT/MultiGhostBench.

作者归属多语言长文本基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。