arXiv:2512.05311cs.LGcs.AI2025-12

研究多次改写后如何区分人类与大模型生成的科学想法

The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?

  • 通过连续改写模拟真实场景,测试模型识别来源的能力
  • 经过五轮改写后,识别准确率平均下降25.4%
  • 越简化、越非专业的表达越难区分来源

随着大模型作为研究代理的广泛应用,区分其与人类生成的想法对理解大模型的研究能力至关重要。尽管文本来源检测已有广泛研究,但科学思想层面的区分仍属空白。本文系统评估了当前最先进的机器学习模型在多次改写后的区分能力。结果显示,经过五轮连续改写后,检测性能平均下降25.4%。此外,引入研究问题作为上下文信息可将检测性能提升最多2.97%。值得注意的是,当思想被改写为简化、非专家风格时,检测算法表现最差,是导致可区分特征消失的主要原因。

原文摘要 · Abstract (English)

With the increasing reliance on LLMs as research agents, distinguishing between LLM and human-generated ideas has become crucial for understanding the cognitive nuances of LLMs' research capabilities. While detecting LLM-generated text has been extensively studied, distinguishing human vs LLM-generated scientific idea remains an unexplored area. In this work, we systematically evaluate the ability of state-of-the-art (SOTA) machine learning models to differentiate between human and LLM-generated ideas, particularly after successive paraphrasing stages. Our findings highlight the challenges SOTA models face in source attribution, with detection performance declining by an average of 25.4\% after five consecutive paraphrasing stages. Additionally, we demonstrate that incorporating the research problem as contextual information improves detection performance by up to 2.97%. Notably, our analysis reveals that detection algorithms struggle significantly when ideas are paraphrased into a simplified, non-expert style, contributing the most to the erosion of distinguishable LLM signatures.

大模型检测思想识别改写攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。