LLM让语言分析既强大又危险,需重构方法以应对伪造文本挑战。
Large Language Models and Forensic Linguistics: Navigating Opportunities and Threats in the Age of Generative AI
- 结合人类判断与AI分析,构建混合型语言鉴定流程。
- 现有检测技术误判率高,对非母语者尤其不公,且易被对抗攻击绕过。
- 适合法律、刑侦领域关注AI文本可信度的研究者参考。
大型语言模型(LLMs)为司法语言学带来双重影响:既是可扩展语料分析与基于嵌入的作者归属分析的强大工具,又因风格模仿、身份隐藏及合成文本泛滥,动摇了个体语言特征(idiolect)的基础假设。近期风格分析研究显示,尽管LLMs能模拟表面语言特征,但与人类写作风格仍存在可检测差异,这对司法实践具有重要含义。然而,当前基于分类器、风格分析或水印的AI文本检测方法普遍存在显著缺陷:对非英语母语者误报率高,且易受同形异义字符等对抗策略影响。这些不确定性使检测结果在法律可采信标准(如Daubert和Kumho Tire框架)下面临质疑。文章建议司法语言学需进行方法论重构,以维持科学可信度与法律可采性,包括采用人机协同工作流、超越二元判断的可解释检测范式,以及在多元群体中评估误差与偏见的验证机制。语言揭示作者信息的核心洞见依然成立,但必须适应日益复杂的人机共同创作链条。
原文摘要 · Abstract (English)
Large language models (LLMs) present a dual challenge for forensic linguistics. They serve as powerful analytical tools enabling scalable corpus analysis and embedding-based authorship attribution, while simultaneously destabilising foundational assumptions about idiolect through style mimicry, authorship obfuscation, and the proliferation of synthetic texts. Recent stylometric research indicates that LLMs can approximate surface stylistic features yet exhibit detectable differences from human writers, a tension with significant forensic implications. However, current AI-text detection techniques, whether classifier-based, stylometric, or watermarking approaches, face substantial limitations: high false positive rates for non-native English writers and vulnerability to adversarial strategies such as homoglyph substitution. These uncertainties raise concerns under legal admissibility standards, particularly the Daubert and Kumho Tire frameworks. The article concludes that forensic linguistics requires methodological reconfiguration to remain scientifically credible and legally admissible. Proposed adaptations include hybrid human-AI workflows, explainable detection paradigms beyond binary classification, and validation regimes measuring error and bias across diverse populations. The discipline's core insight, i.e., that language reveals information about its producer, remains valid but must accommodate increasingly complex chains of human and machine authorship.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。