arXiv:2508.00680cs.CL2025-08被引 2

大模型能精准识别句子级写作风格变化,表现远超传统基准。

Better Call Claude: Can LLMs Detect Changes of Writing Style?

  • 直接用大模型零样本检测句子级风格变化,无需额外训练。
  • 在PAN 2024和2025数据集上准确率显著高于竞赛建议基线。
  • 模型更关注风格本身而非语义内容,适合风格分析研究者使用。

本文研究了当前最先进大语言模型(LLMs)在作者分析中最具挑战性的任务之一——句子级写作风格变化检测中的零样本性能。在官方PAN 2024和2025“多作者写作风格分析”数据集上,对四种LLMs进行基准测试,得出若干观察:第一,最先进的生成式模型对写作风格的细微变化敏感,甚至可识别单个句子的风格差异;第二,其准确率建立了极具挑战性的基准,超越了PAN竞赛所建议的基线方法;第三,我们探讨了语义对模型预测的影响,发现最新一代的LLMs可能比以往报道更敏感于与内容无关的纯风格信号。

原文摘要 · Abstract (English)

This article explores the zero-shot performance of state-of-the-art large language models (LLMs) on one of the most challenging tasks in authorship analysis: sentence-level style change detection. Benchmarking four LLMs on the official PAN~2024 and 2025 "Multi-Author Writing Style Analysis" datasets, we present several observations. First, state-of-the-art generative models are sensitive to variations in writing style - even at the granular level of individual sentences. Second, their accuracy establishes a challenging baseline for the task, outperforming suggested baselines of the PAN competition. Finally, we explore the influence of semantics on model predictions and present evidence suggesting that the latest generation of LLMs may be more sensitive to content-independent and purely stylistic signals than previously reported.

风格检测大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。