评测风格个性化生成效果,发现多指标组合更可靠。
Evaluating Style-Personalized Text Generation: Challenges and Directions
- 提出风格区分基准,测试多种评估指标有效性
- 发现单一指标与人类判断相关性差,集成多指标更优
- 适用于研究个性化文本生成的评估方法设计者
随着大语言模型(LLMs)生成定制化内容能力的提升,风格个性化文本生成——即‘模仿我的写作风格’——已成为快速发展的研究方向。然而,风格个性化具有高度主观性和语境依赖性,极具挑战性。尽管已有研究引入基准和评估指标,但普遍存在非标准化、与人类判断相关性低等问题。现有研究表明,LLMs难以准确捕捉作者特定风格,因此评估指标本身亟需审慎检验。本文系统评估了当前最常用的评估方法,包括BLEU、嵌入向量及以LLM作为评价者等,使用我们提出的风格区分基准进行验证,涵盖三个评估场景下的八项多样化写作任务:领域区分、作者归属和个性化与非个性化文本区分。结果表明,采用多样化的指标集成方法显著优于单一评估方式。文章最后提供可信赖评估风格个性化生成的实践指导。
原文摘要 · Abstract (English)
With the surge of large language models (LLMs) and their ability to produce customized output, style-personalized text generation--"write like me"--has become a rapidly growing area of interest. However, style personalization is highly specific, relative to every user, and depends strongly on the pragmatic context, which makes it uniquely challenging. Although prior research has introduced benchmarks and metrics for this area, they tend to be non-standardized and have known limitations (e.g., poor correlation with human subjects). LLMs have been found to not capture author-specific style well, it follows that the metrics themselves must be scrutinized carefully. In this work we critically examine the effectiveness of the most common metrics used in the field, such as BLEU, embeddings, and LLMs-as-judges. We evaluate these metrics using our proposed style discrimination benchmark, which spans eight diverse writing tasks across three evaluation settings: domain discrimination, authorship attribution, and LLM-generated personalized vs non-personalized discrimination. We find strong evidence that employing ensembles of diverse evaluation metrics consistently outperforms single-evaluator methods, and conclude by providing guidance on how to reliably assess style-personalized text generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。