提出细粒度人机共著文本检测方法,可识别每句话的AI生成比例。
HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring
- 构建自动标注的人机共著文本数据集HACo-Det,支持词级归属标记。
- 微调模型在词级与句级检测中平均F1达0.462以上,优于传统指标法。
- 适用于内容安全、学术诚信等需精确识别生成痕迹的场景。
大语言模型的滥用带来潜在风险,推动了机器生成文本(MGT)检测的发展。现有研究多聚焦二分类、文档级检测,忽略了人类与AI共同创作的文本。本文探索人机共著场景下的细粒度MGT检测可行性。我们提出数据集HACo-Det,通过自动化流程生成带词级归属标签的人机共著文本。将七种主流文档级检测器改造为词级检测器,并在HACo-Det上评估其在词级与句级任务的表现。实验表明,基于指标的方法在细粒度检测中表现不佳,平均F1仅为0.462;而微调模型展现出更优性能与跨领域泛化能力。但研究指出,细粒度共著文本检测仍远未解决,进一步分析了上下文窗口等因素对性能的影响,揭示当前方法局限性,指明改进方向。
原文摘要 · Abstract (English)
The misuse of large language models (LLMs) poses potential risks, motivating the development of machine-generated text (MGT) detection. Existing literature primarily concentrates on binary, document-level detection, thereby neglecting texts that are composed jointly by human and LLM contributions. Hence, this paper explores the possibility of fine-grained MGT detection under human-AI coauthoring. We suggest fine-grained detectors can pave pathways toward coauthored text detection with a numeric AI ratio. Specifically, we propose a dataset, HACo-Det, which produces human-AI coauthored texts via an automatic pipeline with word-level attribution labels. We retrofit seven prevailing document-level detectors to generalize them to word-level detection. Then we evaluate these detectors on HACo-Det on both word- and sentence-level detection tasks. Empirical results show that metric-based methods struggle to conduct fine-grained detection with a 0.462 average F1 score, while finetuned models show superior performance and better generalization across domains. However, we argue that fine-grained co-authored text detection is far from solved. We further analyze factors influencing performance, e.g., context window, and highlight the limitations of current methods, pointing to potential avenues for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。