arXiv:2509.14543cs.CLcs.AI2025-09EMNLP被引 13

大模型难模仿普通人写作的隐性风格,尤其在非正式场景。

Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors

  • 用多种指标评估大模型从少量样本中模仿个人写作风格的能力。
  • 在新闻邮件等正式文本中表现尚可,但在博客论坛等非正式文本中效果差。
  • 提示词数量不足限制个性化效果,适合研究个性化生成的学者参考。

随着大语言模型(LLMs)日益融入个人写作工具,一个关键问题浮现:仅凭少量样本文本,大模型能否忠实模仿个体写作风格?个人风格通常微妙且隐含,难以通过提示明确表达,却对用户对齐生成至关重要。本文对前沿大模型在少样本情境下通过上下文学习模仿个人写作风格的能力进行了全面评估。我们构建了一套互补的评估指标组合,包括作者归属、作者验证、风格匹配与AI检测,以稳健评估风格模仿效果。评估覆盖超过40000次生成,涵盖新闻、邮件、论坛和博客等多领域,涉及400多位真实作者的写作样本。结果显示,尽管大模型在新闻和邮件等结构化文本中能较好逼近用户风格,但在博客和论坛等非正式、细腻的文本中仍表现不佳。进一步分析表明,提示策略(如演示样本数量)存在关键局限。研究揭示了个性化大模型适配的根本差距,亟需改进技术以实现隐性风格一致的生成。为支持未来研究并确保可复现性,我们开源了数据与代码。

原文摘要 · Abstract (English)

As large language models (LLMs) become increasingly integrated into personal writing tools, a critical question arises: can LLMs faithfully imitate an individual's writing style from just a few examples? Personal style is often subtle and implicit, making it difficult to specify through prompts yet essential for user-aligned generation. This work presents a comprehensive evaluation of state-of-the-art LLMs' ability to mimic personal writing styles via in-context learning from a small number of user-authored samples. We introduce an ensemble of complementary metrics-including authorship attribution, authorship verification, style matching, and AI detection-to robustly assess style imitation. Our evaluation spans over 40000 generations per model across domains such as news, email, forums, and blogs, covering writing samples from more than 400 real-world authors. Results show that while LLMs can approximate user styles in structured formats like news and email, they struggle with nuanced, informal writing in blogs and forums. Further analysis on various prompting strategies such as number of demonstrations reveal key limitations in effective personalization. Our findings highlight a fundamental gap in personalized LLM adaptation and the need for improved techniques to support implicit, style-consistent generation. To aid future research and for reproducibility, we open-source our data and code.

风格模仿个性化生成大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。