arXiv:2601.01225cs.CLcs.AI2026-01被引 1

用NLP检测学术文本中人与机器生成内容,识别作者变更和合著者身份。

Stylometry Analysis of Human and Machine Text for Academic Integrity

  • 基于NLP框架,区分人类与机器撰写的学术文本。
  • 在严格指令下生成的文本使检测性能下降,体现对抗性挑战。
  • 适合关注学术诚信、文本溯源与反作弊的研究者使用。

本文针对学术诚信问题,如抄袭、伪造和作者身份验证,提出一种基于自然语言处理(NLP)的框架,通过作者归属和风格变化检测来认证学生内容。相较于已有工作,本研究系统分析了四个任务:(i) 人类与机器文本分类,(ii) 单作者与多作者文本区分,(iii) 多作者文档中的作者变更检测,(iv) 合作撰写文档中的作者识别。所提方法在两个由Gemini生成的数据集上进行评估,采用正常与严格提示两种策略。实验显示,在严格提示生成的数据上性能有所下降,表明精心设计的提示可显著增加机器生成文本的检测难度。相关数据集、代码及其他材料已公开于GitHub,为该领域未来研究提供基准。

原文摘要 · Abstract (English)

This work addresses critical challenges to academic integrity, including plagiarism, fabrication, and verification of authorship of educational content, by proposing a Natural Language Processing (NLP)-based framework for authenticating students' content through author attribution and style change detection. Despite some initial efforts, several aspects of the topic are yet to be explored. In contrast to existing solutions, the paper provides a comprehensive analysis of the topic by targeting four relevant tasks, including (i) classification of human and machine text, (ii) differentiating in single and multi-authored documents, (iii) author change detection within multi-authored documents, and (iv) author recognition in collaboratively produced documents. The solutions proposed for the tasks are evaluated on two datasets generated with Gemini using two different prompts, including a normal and a strict set of instructions. During experiments, some reduction in the performance of the proposed solutions is observed on the dataset generated through the strict prompt, demonstrating the complexities involved in detecting machine-generated text with cleverly crafted prompts. The generated datasets, code, and other relevant materials are made publicly available on GitHub, which are expected to provide a baseline for future research in the domain.

文本溯源学术诚信NLP应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。