arXiv:2604.10101cs.CL2026-04ACL

构建首个中文古诗生成检测基准,评估大模型写诗的可识别性。

Who Wrote This Line? Evaluating the Detection of LLM-Generated Classical Chinese Poetry

论文配图:Who Wrote This Line? Evaluating the Detection of LLM-Generated Classical Chinese Poetry
图 1 · 摘自论文原文
  • 构建含3万首古诗的基准数据集ChangAn,区分人与AI创作。
  • 12种检测器在古诗上准确率普遍低于60%,效果不佳。
  • 适合关注中文生成内容可信度的研究者和平台方。

大语言模型(LLMs)在文学创作领域的快速发展,引发了创意真实性与伦理问题,使检测AI生成文本变得迫切。现有方法主要针对现代汉语,尚未覆盖古典中文诗歌。由于古典诗歌具有严格的格律、共用意象体系及灵活语法,判断其是否由AI生成极具挑战。为此,我们提出ChangAn基准,包含总计30,664首诗歌,其中10,276首为人工创作,20,388首由四款主流大模型生成。基于此,我们系统评估了12种AI检测器在不同文本粒度和生成策略下的表现。结果表明,现有中文文本检测工具在古典诗歌场景下普遍失效,准确率不足60%。该研究验证了ChangAn的有效性与必要性。数据集与代码已开源:https://github.com/VelikayaScarlet/ChangAn。

原文摘要 · Abstract (English)

The rapid development of large language models (LLMs) has extended text generation tasks into the literary domain. However, AI-generated literary creations has raised increasingly prominent issues of creative authenticity and ethics in literary world, making the detection of LLM-generated literary texts essential and urgent. While previous works have made significant progress in detecting AI-generated text, it has yet to address classical Chinese poetry. Due to the unique linguistic features of classical Chinese poetry, such as strict metrical regularity, a shared system of poetic imagery, and flexible syntax, distinguishing whether a poem is authored by AI presents a substantial challenge. To address these issues, we introduce ChangAn, a benchmark for detecting LLM-generated classical Chinese poetry that containing total 30,664 poems, 10,276 are human-written poems and 20,388 poems are generated by four popular LLMs. Based on ChangAn, we conducted a systematic evaluation of 12 AI detectors, investigating their performance variations across different text granularities and generation strategies. Our findings highlight the limitations of current Chinese text detectors, which fail to serve as reliable tools for detecting LLM-generated classical Chinese poetry. These results validate the effectiveness and necessity of our proposed ChangAn benchmark. Our dataset and code are available at https://github.com/VelikayaScarlet/ChangAn.

古诗生成检测基准大模型伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。