用不同人设改写提示词,发现风格差异会显著影响大模型评分。
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
- 用角色化提示模拟多样写作风格,低成本增强评估多样性。
- 相同内容因风格不同,模型表现差距可达显著水平。
- 发现若干通用高低性能触发风格,跨模型均有效。
当前大语言模型(LLM)评估基准往往缺乏写作风格多样性,多遵循标准化格式。这可能导致模型在面对非标准输入时表现脆弱。本文通过基于角色的提示生成方法,低成本地重构评估提示以模拟多样化写作风格。结果显示,即使语义一致,写作风格和提示格式的变化也会显著影响模型性能评估结果。我们识别出若干在多种模型、任务中持续引发高或低表现的独特写作风格,且不依赖模型家族、规模或发布时间。本工作提供了一种可扩展的方法,用于增强现有基准,提升其在语言变体下的外部有效性。
原文摘要 · Abstract (English)
Current benchmarks for evaluating Large Language Models (LLMs) often do not exhibit enough writing style diversity, with many adhering primarily to standardized conventions. Such benchmarks do not fully capture the rich variety of communication patterns exhibited by humans. Thus, it is possible that LLMs, which are optimized on these benchmarks, may demonstrate brittle performance when faced with "non-standard" input. In this work, we test this hypothesis by rewriting evaluation prompts using persona-based LLM prompting, a low-cost method to emulate diverse writing styles. Our results show that, even with identical semantic content, variations in writing style and prompt formatting significantly impact the estimated performance of the LLM under evaluation. Notably, we identify distinct writing styles that consistently trigger either low or high performance across a range of models and tasks, irrespective of model family, size, and recency. Our work offers a scalable approach to augment existing benchmarks, improving the external validity of the assessments they provide for measuring LLM performance across linguistic variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。