同时检测和溯源大模型生成文本,提升多语言场景下的识别能力。
Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- 采用多任务学习框架,联合优化文本检测与作者溯源。
- 在九个数据集上表现优异,支持多语言和多种大模型源。
- 可抵御对抗性混淆,适合安全审计与内容溯源场景。
大型语言模型(如 GPT-4、Llama)在生成自然语言方面表现出色,但也带来安全与可信度挑战。现有方法主要针对英语文本的生成内容检测,而对具体模型来源的作者溯源研究较少。本文提出 DA-MTL 框架,通过多任务学习同时实现文本检测与作者溯源。我们在九个数据集和四种主干模型上进行评估,验证了其在多语言和多种模型来源下的强性能。该框架能捕捉各任务特性并共享知识,提升双重任务表现。此外,我们分析了跨模态与跨语言模式,并测试了对对抗性混淆的鲁棒性。结果为理解 LLM 行为及检测与溯源的泛化能力提供了重要洞见。
原文摘要 · Abstract (English)
Large Language Models (LLMs), such as GPT-4 and Llama, have demonstrated remarkable abilities in generating natural language. However, they also pose security and integrity challenges. Existing countermeasures primarily focus on distinguishing AI-generated content from human-written text, with most solutions tailored for English. Meanwhile, authorship attribution--determining which specific LLM produced a given text--has received comparatively little attention despite its importance in forensic analysis. In this paper, we present DA-MTL, a multi-task learning framework that simultaneously addresses both text detection and authorship attribution. We evaluate DA-MTL on nine datasets and four backbone models, demonstrating its strong performance across multiple languages and LLM sources. Our framework captures each task's unique characteristics and shares insights between them, which boosts performance in both tasks. Additionally, we conduct a thorough analysis of cross-modal and cross-lingual patterns and assess the framework's robustness against adversarial obfuscation techniques. Our findings offer valuable insights into LLM behavior and the generalization of both detection and authorship attribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。