用大模型对文本风格迁移能力打分,无需标注即可验证作者身份
One-shot Style Transfer LLM log-probabilities for Authorship Attribution and Verification
- 基于大模型的对数概率计算文本间风格转移程度
- 在作者验证任务中性能超越提示工程基线,且随模型规模提升持续增强
- 支持多语言,可灵活权衡计算开销与准确率
计算风格学通过量化文本模式研究写作风格,可用于作者身份识别、身份关联和抄袭检测。尽管语言建模与此类任务密切相关,现代大语言模型(LLMs)的预训练却未被充分应用于作者身份验证。本文提出一种无监督框架,利用LLM的对数概率衡量两段文本间的风格转移能力。该方法充分利用了因果语言建模(CLM)预训练、单次输入能力和模型规模,避免显式标注。实验表明,在相似模型规模下,该方法显著优于基于提示的无监督基线,在足够规模模型下,性能可媲美或超越对比学习基线。跨语言测试也表现出色。框架效果随模型规模持续提升。针对作者验证任务,我们还设计了一种增加测试时计算量的机制,实现计算成本与性能之间的灵活权衡。
原文摘要 · Abstract (English)
Computational stylometry studies writing style through quantitative textual patterns, enabling applications such as authorship attribution, identity linking, and plagiarism detection. Despite the relevance of language modeling to these tasks, the pre-training of modern large language models (LLMs) has been underutilized in authorship attribution and verification. We introduce an unsupervised framework that uses the log-probabilities of an LLM to measure style transferability between two texts. This framework takes advantage of the extensive Causal Language Modeling (CLM) pre-training, one-shot capabilities and scale of LLMs, avoiding explicit supervision. Our methods substantially outperform prompting-based unsupervised baselines in authorship verification at similar model sizes, and is competitive with or improves contrastive baselines in most settings with sufficient model scale. We further observe strong performance across non-English languages. The effectiveness of the proposed framework improves consistently with increasing model scale. In the case of authorship verification, we propose an additional mechanism that increases test-time computation to improve accuracy; enabling flexible trade-offs between computational cost and task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。