论文发现学术导师与学生写作风格相似,可能误导作者归属判断。
Writing Style Similarity Reflects Academic Genealogy
- 用数学家谱数据构建作者语料库,验证师生风格关联性。
- 导师与学生风格相似度比同领域随机作者高39.9%。
- 学术兄弟(同导师不同校)也显著相似,适合研究学术传承者。
随着作者归属系统被用于检测代笔和AI生成论文,其错误可能误伤合法作者。这类系统常将写作风格相似等同于个人身份。但研究人员在导师指导下学习,会继承其写作风格特征。本文基于数学家谱项目图谱,构建了包含5,803位拥有至少两篇独立论文的arXiv作者的语料库,包含2,501对真实导师-学生关系。通过微调模型的嵌入表示,导师与学生的余弦距离比同领域随机作者近39.9%。使用两个开源模型复现该效应,分别达到12.6%和14.5%。学术兄弟(同一导师的学生,可能从未见面)在8,360对中平均接近30.4%,即使就读于不同机构。仅共享机构与领域的作者对则无显著相似性。在一个封闭集归属任务中,系统错误发生于真实作者的导师、学生、学术兄弟或实验室同伴的概率,是随机情况的11倍。
原文摘要 · Abstract (English)
As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusations against legitimate authors. These systems conflate stylistic similarity with individual identity. Researchers, however, study under advisors, and inherit their stylistic quirks. We build a corpus of arXiv authors with $\geq 2$ solo papers from the Mathematics Genealogy Project graph, giving $5{,}803$ total authors and $2{,}501$ ground-truth advisor-student pairings. Using embeddings from a fine-tuned model, advisors sit $39.9\%$ closer in cosine distance to their students than a random same-field author does. Using two open models, we reproduce the effect at $12.6\%$ and $14.5\%$. Academic siblings, two students of one advisor who may never have met, sit $30.4\%$ closer across $8{,}360$ pairs, even when they studied at different institutions. Pairs who share only institution and field show negligible similarity. Given a closed-set attribution task over the same corpus, the system's errors occur on the true author's advisor, student, academic sibling, or lab mate $11$ times more often than chance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。