通过分析推理图结构,提升大模型文本作者归属的鲁棒性。
Show Me How You Reason and I'll Tell You Who You Are: Reasoning Graphs for Robust LLM Authorship Attribution

- 用论点挖掘提取文本推理图,构建图神经网络进行分析。
- 在改写和反向翻译攻击下,准确率比基线高27个百分点。
- 对未见模型版本生成文本仍有效,适合真实场景应用。
随着大语言模型(LLMs)在各类场景中的广泛应用,检测其生成文本及识别作者身份已成为紧迫问题。以往研究主要依赖表层语言特征,易受改写等混淆技术影响。本文超越语言表面,通过论点挖掘管道提取并分析大模型生成文本中的推理结构,以捕捉更复杂的模型作者特征。提出一种基于推理图的图神经网络方法,在与传统Longformer基线对比中表现更优。在改写和反向翻译等混淆攻击下,准确率最高提升27个百分点;在未见模型版本生成文本上,准确率提升19个百分点,模拟了新模型持续发布的真实场景。
原文摘要 · Abstract (English)
Given the current trend to employ large language models (LLMs) in almost any imaginable context, LLM-generated text detection and authorship attribution have become a pressing issue. Prior work has primarily focused on surface-level linguistic features, an approach shown to be susceptible to paraphrasing and other obfuscation techniques. In this paper, we go beyond the linguistic surface, extracting and analysing reasoning structures in LLM-generated texts with the goal of capturing more complex signals of LLM authorship. We propose a graph neural network approach that leverages reasoning graphs extracted by an argument mining pipeline, demonstrating improved robustness and generalisation over a traditional Longformer baseline. Our approach outperforms the baseline by up to 27 percentage points under the obfuscation attacks such as paraphrasing and backtranslation, and 19 percentage points when evaluated on the texts generated by the unseen model versions, simulating real-world conditions in which new LLM versions are continuously released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。