区分文本中作者与编辑角色,精准识别LLM生成文本类型
Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection
- 用修辞结构理论构建创作者逻辑图,提取编辑风格特征
- 在四分类任务中超越12个基线模型,误报率低
- 适合需要精细治理的AI内容监管场景
大语言模型(LLMs)的滥用亟需对合成文本进行精准检测。现有方法多采用二分类或三分类设置,仅能区分纯人工或纯LLM文本,或最多识别协作文本,难以满足精细化治理需求,因为经LLM润色的人类文本与拟人化LLM文本可能触发不同政策后果。本文在严格的四分类设定下探索细粒度的LLM生成文本检测。为此,提出RACE(修辞分析用于创作者-编辑建模),通过修辞结构理论(RST)构建创作者的基础逻辑图,并提取编辑在基本话语单元(EDU)层面的风格特征。实验表明,RACE在识别细粒度类型上优于12个基线模型,且误报率低,为LLM治理提供了符合政策意图的解决方案。
原文摘要 · Abstract (English)
The misuse of large language models (LLMs) requires precise detection of synthetic text. Existing works mainly follow binary or ternary classification settings, which can only distinguish pure human/LLM text or collaborative text at best. This remains insufficient for the nuanced regulation, as the LLM-polished human text and humanized LLM text often trigger different policy consequences. In this paper, we explore fine-grained LLM-generated text detection under a rigorous four-class setting. To handle such complexities, we propose RACE (Rhetorical Analysis for Creator-Editor Modeling), a fine-grained detection method that characterizes the distinct signatures of creator and editor. Specifically, RACE utilizes Rhetorical Structure Theory (RST) to construct a logic graph for the creator's foundation while extracting Elementary Discourse Unit (EDU)-level features for the editor's style. Experiments show that RACE outperforms 12 baselines in identifying fine-grained types with low false alarms, offering a policy-aligned solution for LLM regulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。