arXiv:2512.04838cs.CL2025-12Conference of the …被引 3

通过可解释标注定位人机混合文本的生成边界。

DAMASHA: Detecting AI in Mixed Adversarial Texts via Segmentation with Human-interpretable Attribution

  • 融合风格特征与困惑度信号,建模文本生成边界。
  • 在对抗性攻击下仍保持高分割准确率,优于现有方法。
  • 提供人类可读的归因图,助于理解模型决策过程。

随着大语言模型的发展,人机生成文本的界限日益模糊。本文提出Info-Mask框架,用于识别人机混合文本中的作者转换点,这对真实性、可信度和人工监管具有重要意义。该框架结合风格特征、困惑度驱动信号和结构化边界建模,实现对协作式人机内容的精准分割。为评估系统鲁棒性,我们构建并发布对抗性基准数据集Mixed-text Adversarial setting for Segmentation (MAS),用于测试现有检测器的极限。除分割精度外,本文引入人类可解释归因(HIA)叠加层,展示风格特征如何影响边界预测,并开展小规模人工研究验证其有效性。在多种架构下,Info-Mask显著提升对抗条件下的段级鲁棒性,建立新基准,同时揭示现存挑战。研究结果凸显了可对抗攻击、可解释的人机共著检测的潜力与局限,对人机协作中的信任与监督具有重要启示。

原文摘要 · Abstract (English)

In the age of advanced large language models (LLMs), the boundaries between human and AI-generated text are becoming increasingly blurred. We address the challenge of segmenting mixed-authorship text, that is identifying transition points in text where authorship shifts from human to AI or vice-versa, a problem with critical implications for authenticity, trust, and human oversight. We introduce a novel framework, called Info-Mask for mixed authorship detection that integrates stylometric cues, perplexity-driven signals, and structured boundary modeling to accurately segment collaborative human-AI content. To evaluate the robustness of our system against adversarial perturbations, we construct and release an adversarial benchmark dataset Mixed-text Adversarial setting for Segmentation (MAS), designed to probe the limits of existing detectors. Beyond segmentation accuracy, we introduce Human-Interpretable Attribution (HIA overlays that highlight how stylometric features inform boundary predictions, and we conduct a small-scale human study assessing their usefulness. Across multiple architectures, Info-Mask significantly improves span-level robustness under adversarial conditions, establishing new baselines while revealing remaining challenges. Our findings highlight both the promise and limitations of adversarially robust, interpretable mixed-authorship detection, with implications for trust and oversight in human-AI co-authorship.

文本检测可解释性对抗防御人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。