arXiv:2510.00890cs.CLcs.AI2025-10被引 14

精准定位论文中AI生成段落,提升跨领域检测可靠性。

Span-level Detection of AI-generated Scientific Text via Contrastive Learning and Structural Calibration

  • 结合章节风格建模与多级对比学习,捕捉人机写作细微差异。
  • 在10万样本数据集上实现AI段落检测F1达80.17,跨度检测F1为74.36。
  • 支持细粒度定位与置信度校准,适合学术界与出版机构使用。

大型语言模型在科学写作中的广泛应用引发了作者身份真实性与学术出版可靠性的严重担忧。现有检测方法多依赖文档级分类或表面统计特征,忽视细粒度段落定位,校准能力弱,且跨学科泛化性差。为此,我们提出Sci-SpanDet,一种结构感知的AI生成学术文本检测框架。该方法结合章节条件风格建模与多层级对比学习,捕捉细微的人机写作风格差异,缓解主题依赖,提升跨领域鲁棒性;同时融合BIO-CRF序列标注与基于指针的边界解码及置信度校准,实现精确的段落级检测与可靠的概率估计。在新构建的跨学科数据集(包含来自GPT、Qwen、DeepSeek、LLaMA等多类模型生成的10万条标注样本)上,Sci-SpanDet表现优异,AI检测F1达80.17,AUROC为92.63,段落级F1为74.36。此外,其对对抗改写具有强鲁棒性,且在IMRaD各部分与不同学科间保持均衡准确率,显著优于现有基线。为促进可复现研究,数据集与源码将在发表后公开。

原文摘要 · Abstract (English)

The rapid adoption of large language models (LLMs) in scientific writing raises serious concerns regarding authorship integrity and the reliability of scholarly publications. Existing detection approaches mainly rely on document-level classification or surface-level statistical cues; however, they neglect fine-grained span localization, exhibit weak calibration, and often fail to generalize across disciplines and generators. To address these limitations, we present Sci-SpanDet, a structure-aware framework for detecting AI-generated scholarly texts. The proposed method combines section-conditioned stylistic modeling with multi-level contrastive learning to capture nuanced human-AI differences while mitigating topic dependence, thereby enhancing cross-domain robustness. In addition, it integrates BIO-CRF sequence labeling with pointer-based boundary decoding and confidence calibration to enable precise span-level detection and reliable probability estimates. Extensive experiments on a newly constructed cross-disciplinary dataset of 100,000 annotated samples generated by multiple LLM families (GPT, Qwen, DeepSeek, LLaMA) demonstrate that Sci-SpanDet achieves state-of-the-art performance, with F1(AI) of 80.17, AUROC of 92.63, and Span-F1 of 74.36. Furthermore, it shows strong resilience under adversarial rewriting and maintains balanced accuracy across IMRaD sections and diverse disciplines, substantially surpassing existing baselines. To ensure reproducibility and to foster further research on AI-generated text detection in scholarly documents, the curated dataset and source code will be publicly released upon publication.

AI检测学术写作细粒度定位对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。