arXiv:2606.06481cs.CLcs.AI2026-06被引 1

构建渐进式人机协作文本演化基准,揭示AI写作痕迹的动态变化规律。

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

论文配图:Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection
图 1 · 摘自论文原文
  • 按预设编辑操作与覆盖率,生成九版逐步演变的文本
  • 混合作者中间版本检测难度高于纯人工或纯AI版本
  • 适用于评估多粒度检测模型在真实协作场景下的表现

随着AI写作助手深度融入实际撰写与修改流程,文档已不再是纯粹的人类创作或AI生成,而是由渐进式人机协同编辑形成。现有AI文本检测基准大多仅关注最终输出,难以理解AI署名信号如何在修订过程中产生、累积或消失。我们提出OpAI-Bench,一个基于操作引导的基准,用于研究文档、句子、词元和片段四个粒度下的人类到AI文本转化过程。从人类撰写文档出发,在预设的五种代表性AI编辑操作和九个不同AI覆盖水平下,构建每个样本的九个连续修订版本,覆盖四个领域,并在多粒度上保留完整的作者溯源信息。该基准支持8个文档级、7个句子级和2个细粒度(词元/片段)检测器的全面评估。实验表明,AI文本可检测性不仅取决于AI编辑内容比例,还受编辑操作、领域和累积修订历史影响。有趣的是,混合作者的中间版本往往比完全人工或高度AI化的终点版本更难检测,暴露出现有基准忽视的非单调检测模式。OpAI-Bench为分析在真实渐进编辑场景中AI辅助写作何时、如何以及是否可被检测提供了可控测试平台。代码与数据集开源:https://github.com/VILA-Lab/OpAI-Bench。

原文摘要 · Abstract (English)

As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-written or AI-generated, but instead result from progressive human-AI co-editing. However, existing AI-text detection benchmarks largely focus on final outputs and provide limited understanding of how AI authorship signals emerge, accumulate, or disappear throughout the revision process. We introduce OpAI-Bench, an operation-guided benchmark for studying progressive human-to-AI text transformation across document, sentence, token, and span granularities. Starting from human-written documents, OpAI-Bench constructs nine sequentially revised versions for each sample under predefined AI coverage levels and five representative AI edit operations, covering four domains while preserving complete authorship provenance at multiple granularities. The benchmark supports comprehensive evaluation with 8 document-level detectors, 7 sentence-level detectors, and 2 fine-grained token/span-level detectors. Experiments reveal that AI-text detectability is governed not only by the proportion of AI-edited content, but also by edit operation, domain, and cumulative revision history. Interestingly, we notice that mixed-authorship intermediate versions are often harder to detect than both fully human and heavily AI-edited endpoints, exposing non-monotonic detection patterns missed by existing benchmarks. OpAI-Bench provides a controlled testbed for analyzing whether, when, and how AI-assisted writing becomes detectable under realistic progressive editing scenarios. Our code and benchmark are available at https://github.com/VILA-Lab/OpAI-Bench.

AI检测人机协作文本演化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。