arXiv:2602.01560cs.CL2026-02被引 1

用罕见度评估作文原创性,发现AI作文质量高但想法雷同。

Argument Rarity-based Originality Assessment for AI-Assisted Writing

  • 以语料库中论点的罕见程度定义原创性,从结构、主张、证据和认知深度四方面量化。
  • 人类作文主张罕见度是AI的4.6倍,且质量越高主张越不罕见,存在质量与原创性权衡。
  • 四个维度独立性强,适合用于区分人工写作与AI生成内容的深层差异。

本文提出一种名为论证罕见度原创性评估(AROA)的框架,用于自动评估学生作文中的论点原创性。AROA将原创性定义为在参考语料库中的罕见程度,并通过结构罕见度、主张罕见度、证据罕见度和认知深度四个互补组件进行评估,采用密度估计方法量化并结合质量调整。基于1,375篇人工作文和1,000篇AI生成作文在两个论题上的实验显示:第一,文本质量与主张罕见度呈强负相关(r = -0.67),表明存在质量-原创性权衡;第二,尽管AI作文质量接近完美(Q = 0.998),其主张罕见度仅为人类水平的约五分之一(AI: 0.037,human: 0.170),说明大模型可复现结构但缺乏语义原创性;第三,四个维度间相关性极低(结构与语义维度间r = 0.06–0.13),证明它们捕捉的是原创性的独立维度。结果表明,在人工智能时代,写作评价应从关注质量转向重视原创性。

原文摘要 · Abstract (English)

This study proposes Argument Rarity-based Originality Assessment (AROA), a framework for automatically evaluating argumentative originality in student essays. AROA defines originality as rarity within a reference corpus and evaluates it through four complementary components: structural rarity, claim rarity, evidence rarity, and cognitive depth, quantified via density estimation and integrated with quality adjustment. Experiments using 1,375 human essays and 1,000 AI-generated essays on two argumentative topics revealed three key findings. First, a strong negative correlation (r = -0.67) between text quality and claim rarity demonstrates a quality-originality trade-off. Second, while AI essays achieved near-perfect quality scores (Q = 0.998), their claim rarity was approximately one-fifth of human levels (AI: 0.037, human: 0.170), indicating that LLMs can reproduce argumentative structure but not semantic originality. Third, the four components showed low mutual correlations (r = 0.06--0.13 between structural and semantic dimensions), confirming that they capture genuinely independent aspects of originality. These results suggest that writing assessment in the AI era must shift from quality to originality.

原创性评估AI写作教育评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。