构建首个中文现代诗生成检测基准,揭示现有工具失效
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
- 自建800首真人+4.16万首AI诗数据集,覆盖六位诗人与四款主流模型
- 六种检测器在该数据集上表现均不理想,尤其难以识别风格特征
- 为中文诗歌生态提供关键检测工具基础,适合关注AI内容治理者
大型语言模型的快速发展使生成文本与人类写作几乎无法区分。已有文本检测研究虽取得进展,但尚未涉及现代中文诗歌。由于现代中文诗歌的独特性,判断其是否由人工智能生成极具挑战。随着AI生成现代诗泛滥,严重扰乱诗歌生态。针对中文世界中识别此类诗歌的紧迫需求,本文提出首个专门用于检测LLM生成现代中文诗歌的基准。首先构建高质量数据集,包含6位专业诗人创作的800首诗,以及4款主流大模型生成的41,600首诗。随后对6种检测器在此数据集上进行系统评估。实验结果表明,当前检测器无法作为可靠工具识别大模型生成的现代诗。最难检测的是内在特质,尤其是诗歌风格。检测结果验证了所提基准的有效性与必要性。本工作为未来AI生成诗歌检测奠定基础。
原文摘要 · Abstract (English)
The rapid development of advanced large language models (LLMs) has made AI-generated text indistinguishable from human-written text. Previous work on detecting AI-generated text has made effective progress, but has not involved modern Chinese poetry. Due to the distinctive characteristics of modern Chinese poetry, it is difficult to identify whether a poem originated from humans or AI. The proliferation of AI-generated modern Chinese poetry has significantly disrupted the poetry ecosystem. Based on the urgency of identifying AI-generated poetry in the real Chinese world, this paper proposes a novel benchmark for detecting LLMs-generated modern Chinese poetry. We first construct a high-quality dataset, which includes both 800 poems written by six professional poets and 41,600 poems generated by four mainstream LLMs. Subsequently, we conduct systematic performance assessments of six detectors on this dataset. Experimental results demonstrate that current detectors cannot be used as reliable tools to detect modern Chinese poems generated by LLMs. The most difficult poetic features to detect are intrinsic qualities, especially style. The detection results verify the effectiveness and necessity of our proposed benchmark. Our work lays a foundation for future detection of AI-generated poetry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。