综述作者分析的最新进展与挑战,覆盖机器学习到大模型的应用。
Trends and Challenges in Authorship Analysis: A Review of ML, DL, and LLM Approaches
- 系统梳理2015-2024年作者归属与验证的主流方法演进
- 揭示低资源语言、跨域泛化等关键研究空白
- 适合想了解该领域前沿与痛点的研究者参考
作者分析在法语语言学、学术界、网络安全和数字内容认证等领域具有重要意义。本文对2015至2024年间作者归属与作者验证两个核心子任务进行系统文献综述,涵盖传统机器学习、深度学习模型及大语言模型的最新方法,分析其技术演进、优势与局限。研究总结了各类方法所采用的特征提取技术、使用数据集,并指出当前面临的关键挑战:低资源语言处理、多语言适应性、跨域泛化能力以及对AI生成文本的检测。本综述旨在为研究人员提供领域最新趋势与挑战的全景图,推动更可靠、准确的作者分析系统在多样化文本场景中的发展。
原文摘要 · Abstract (English)
Authorship analysis plays an important role in diverse domains, including forensic linguistics, academia, cybersecurity, and digital content authentication. This paper presents a systematic literature review on two key sub-tasks of authorship analysis; Author Attribution and Author Verification. The review explores SOTA methodologies, ranging from traditional ML approaches to DL models and LLMs, highlighting their evolution, strengths, and limitations, based on studies conducted from 2015 to 2024. Key contributions include a comprehensive analysis of methods, techniques, their corresponding feature extraction techniques, datasets used, and emerging challenges in authorship analysis. The study highlights critical research gaps, particularly in low-resource language processing, multilingual adaptation, cross-domain generalization, and AI-generated text detection. This review aims to help researchers by giving an overview of the latest trends and challenges in authorship analysis. It also points out possible areas for future study. The goal is to support the development of better, more reliable, and accurate authorship analysis system in diverse textual domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。